We gratefully acknowledge support from
the Simons Foundation and member institutions.
Full-text links:

Download:

Current browse context:

math.OC

Change to browse by:

References & Citations

Bookmark

(what is this?)
CiteULike logo BibSonomy logo Mendeley logo del.icio.us logo Digg logo Reddit logo

Mathematics > Optimization and Control

Title: Unified Convergence Analysis of Stochastic Momentum Methods for Convex and Non-convex Optimization

Abstract: Recently, stochastic momentum methods have been widely adopted in training deep neural networks. However, their convergence analysis is still underexplored at the moment, in particular for non-convex optimization. This paper fills the gap between practice and theory by developing a basic convergence analysis of two stochastic momentum methods, namely stochastic heavy-ball method and the stochastic variant of Nesterov's accelerated gradient method. We hope that the basic convergence results developed in this paper can serve the reference to the convergence of stochastic momentum methods and also serve the baselines for comparison in future development. The novelty of convergence analysis presented in this paper is a unified framework, revealing more insights about the similarities and differences between different stochastic momentum methods and stochastic gradient method. The unified framework also exhibits a continuous change from the gradient method to Nesterov's accelerated gradient method and the heavy-ball method incurred by a free parameter. The theoretical and empirical results show that the stochastic variant of Nesterov's accelerated gradient method achieves a good tradeoff for optimizing deep neural networks among the three stochastic methods.
Subjects: Optimization and Control (math.OC); Machine Learning (stat.ML)
Cite as: arXiv:1604.03257 [math.OC]
  (or arXiv:1604.03257v1 [math.OC] for this version)

Submission history

From: Tianbao Yang [view email]
[v1] Tue, 12 Apr 2016 06:24:19 GMT (51kb)
[v2] Wed, 4 May 2016 23:11:39 GMT (146kb)

Link back to: arXiv, form interface, contact.