Lossy Gradient Compression: How Much Accuracy Can One Bit Buy?

Salehkalaibar, Sadaf; Rini, Stefano

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 2202

Computer Science > Machine Learning

Title: Lossy Gradient Compression: How Much Accuracy Can One Bit Buy?

Authors: Sadaf Salehkalaibar, Stefano Rini

(Submitted on 6 Feb 2022 (this version), latest version 3 Jun 2022 (v2))

Abstract: In federated learning (FL), a global model is trained at a Parameter Server (PS) by aggregating model updates obtained from multiple remote learners. Critically, the communication between the remote users and the PS is limited by the available power for transmission, while the transmission from the PS to the remote users can be considered unbounded. This gives rise to the distributed learning scenario in which the updates from the remote learners have to be compressed so as to meet communication rate constraints in the uplink transmission toward the PS. For this problem, one would like to compress the model updates so as to minimize the resulting loss in accuracy. In this paper, we take a rate-distortion approach to answer this question for the distributed training of a deep neural network (DNN). In particular, we define a measure of the compression performance, the \emph{per-bit accuracy}, which addresses the ultimate model accuracy that a bit of communication brings to the centralized model. In order to maximize the per-bit accuracy, we consider modeling the gradient updates at remote learners as a generalized normal distribution. Under this assumption on the model update distribution, we propose a class of distortion measures for the design of quantizer for the compression of the model updates. We argue that this family of distortion measures, which we refer to as "$M$-magnitude weighted $L_2$" norm, capture the practitioner intuition in the choice of gradient compressor. Numerical simulations are provided to validate the proposed approach.

Subjects:	Machine Learning (cs.LG); Information Theory (cs.IT)
Cite as:	arXiv:2202.02812 [cs.LG]
	(or arXiv:2202.02812v1 [cs.LG] for this version)

Submission history

From: Sadaf Salehkalaibar Dr [view email]
[v1] Sun, 6 Feb 2022 16:29:00 GMT (98kb)
[v2] Fri, 3 Jun 2022 03:09:46 GMT (112kb)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2202.02812v1

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: Lossy Gradient Compression: How Much Accuracy Can One Bit Buy?

Submission history