Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration

Chen, Congliang; Shen, Li; Zou, Fangyu; Liu, Wei

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 2101

Computer Science > Machine Learning

Title: Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration

Authors: Congliang Chen, Li Shen, Fangyu Zou, Wei Liu

(Submitted on 14 Jan 2021 (v1), last revised 8 Aug 2022 (this version, v2))

Abstract: Adam is one of the most influential adaptive stochastic algorithms for training deep neural networks, which has been pointed out to be divergent even in the simple convex setting via a few simple counterexamples. Many attempts, such as decreasing an adaptive learning rate, adopting a big batch size, incorporating a temporal decorrelation technique, seeking an analogous surrogate, \textit{etc.}, have been tried to promote Adam-type algorithms to converge. In contrast with existing approaches, we introduce an alternative easy-to-check sufficient condition, which merely depends on the parameters of the base learning rate and combinations of historical second-order moments, to guarantee the global convergence of generic Adam for solving large-scale non-convex stochastic optimization. This observation, coupled with this sufficient condition, gives much deeper interpretations on the divergence of Adam. On the other hand, in practice, mini-Adam and distributed-Adam are widely used without any theoretical guarantee. We further give an analysis on how the batch size or the number of nodes in the distributed system affects the convergence of Adam, which theoretically shows that mini-batch and distributed Adam can be linearly accelerated by using a larger mini-batch size or a larger number of nodes.At last, we apply the generic Adam and mini-batch Adam with the sufficient condition for solving the counterexample and training several neural networks on various real-world datasets. Experimental results are exactly in accord with our theoretical analysis.

Comments:	Accepted to JMLR(JMLR). arXiv admin note: substantial text overlap with arXiv:1811.09358
Subjects:	Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC); Optimization and Control (math.OC)
Cite as:	arXiv:2101.05471 [cs.LG]
	(or arXiv:2101.05471v2 [cs.LG] for this version)

Submission history

From: Li Shen [view email]
[v1] Thu, 14 Jan 2021 06:42:29 GMT (817kb,D)
[v2] Mon, 8 Aug 2022 07:25:27 GMT (1210kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2101.05471

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: Towards Practical Adam: Non-Convexity, Convergence Theory, and Mini-Batch Acceleration

Submission history