Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology

Nguyen, Quynh; Mondelli, Marco

Full-text links:

Download:

Current browse context:

cs.LG

< prev | next >

new | recent | 2002

Computer Science > Machine Learning

Title: Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology

Authors: Quynh Nguyen, Marco Mondelli

(Submitted on 18 Feb 2020 (v1), last revised 17 Dec 2020 (this version, v3))

Abstract: Recent works have shown that gradient descent can find a global minimum for over-parameterized neural networks where the widths of all the hidden layers scale polynomially with $N$ ($N$ being the number of training samples). In this paper, we prove that, for deep networks, a single layer of width $N$ following the input layer suffices to ensure a similar guarantee. In particular, all the remaining layers are allowed to have constant widths, and form a pyramidal topology. We show an application of our result to the widely used LeCun's initialization and obtain an over-parameterization requirement for the single wide layer of order $N^2.$

Comments:	Accepted at NeurIPS 2020
Subjects:	Machine Learning (cs.LG); Machine Learning (stat.ML)
Cite as:	arXiv:2002.07867 [cs.LG]
	(or arXiv:2002.07867v3 [cs.LG] for this version)

Submission history

From: Quynh Nguyen [view email]
[v1] Tue, 18 Feb 2020 20:21:27 GMT (104kb,D)
[v2] Mon, 1 Jun 2020 15:23:02 GMT (86kb,D)
[v3] Thu, 17 Dec 2020 19:45:04 GMT (124kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2002.07867

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Machine Learning

Title: Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology

Submission history