A mean field view of the landscape of two-layer neural networks.

A mean field view of the landscape of two-layer neural networks.
复制标题

DOI:
10.1073/pnas.1806579115
复制
发表时间:
2018-08-14
影响因子:
11.1
通讯作者:
Nguyen PM
Nguyen PM
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Mei S;Montanari A;Nguyen PM

文献摘要

参考文献

被引文献

相似文献

事实证明,多层神经网络在从图像分类到机器人等各种任务中都非常成功。然而,这种实际成功的原因及其确切的适用范围尚不清楚。从数据中学习神经网络需要解决具有数百万变量的复杂优化问题。这是通过随机梯度下降(SGD)算法完成的。我们研究的情况下,两层网络,并得出一个紧凑的描述SGD动力学的限制偏微分方程。除其他后果外,这表明SGD动态不会随着网络规模的增加而变得更加复杂。多层神经网络是机器学习中最强大的模型之一,但这种成功的根本原因却无法从数学上理解。学习神经网络需要优化一个非凸的高维目标(风险函数),这是一个通常使用随机梯度下降(SGD)来解决的问题。SGD收敛于风险的全局最优值还是局部最优值?在前一种情况下,发生这种情况是因为局部极小值不存在,还是因为SGD以某种方式避免了它们?在后者中,为什么SGD所达到的局部极小值具有良好的推广特性?在本文中,我们考虑一个简单的情况下,即两层神经网络,并证明,在一个适当的缩放限制SGD动态捕获的一定的非线性偏微分方程(PDE),我们称之为分布动力学(DD)。然后,我们考虑几个具体的例子,并显示如何DD可以用来证明收敛的SGD网络与近理想的泛化误差。这种描述允许“平均”神经网络景观的一些复杂性,并可用于证明噪声SGD的一般收敛结果。
Multilayer neural networks have proven extremely successful in a variety of tasks, from image classification to robotics. However, the reasons for this practical success and its precise domain of applicability are unknown. Learning a neural network from data requires solving a complex optimization problem with millions of variables. This is done by stochastic gradient descent (SGD) algorithms. We study the case of two-layer networks and derive a compact description of the SGD dynamics in terms of a limiting partial differential equation. Among other consequences, this shows that SGD dynamics does not become more complex when the network size increases. Multilayer neural networks are among the most powerful models in machine learning, yet the fundamental reasons for this success defy mathematical understanding. Learning a neural network requires optimizing a nonconvex high-dimensional objective (risk function), a problem that is usually attacked using stochastic gradient descent (SGD). Does SGD converge to a global optimum of the risk or only to a local optimum? In the former case, does this happen because local minima are absent or because SGD somehow avoids them? In the latter, why do local minima reached by SGD have good generalization properties? In this paper, we consider a simple case, namely two-layer neural networks, and prove that—in a suitable scaling limit—SGD dynamics is captured by a certain nonlinear partial differential equation (PDE) that we call distributional dynamics (DD). We then consider several specific examples and show how DD can be used to prove convergence of SGD to networks with nearly ideal generalization error. This description allows for “averaging out” some of the complexities of the landscape of neural networks and can be used to prove a general convergence result for noisy SGD.
DOI: 10.1109/18.556601
发表时间: 1996-11-01
影响因子: 2.5
作者:
Lee, WS;Bartlett, PL;Williamson, RC
通讯作者: Williamson, RC
DOI: 10.1137/s0036141096303359
发表时间: 1998-01-01
影响因子: 2
作者:
Jordan, R;Kinderlehrer, D;Otto, F
通讯作者: Otto, F
DOI: 10.1109/18.661502
发表时间: 1998-03-01
影响因子: 2.5
作者:
Bartlett, PL
通讯作者: Bartlett, PL
DOI: 10.1088/0953-8984/11/10a/011
发表时间: 1999-03-15
影响因子: 2.7
作者:
Mézard, M;Parisi, G
通讯作者: Parisi, G
DOI: 10.1145/3065386
发表时间: 2017-06-01
影响因子: 22.7
作者:
Krizhevsky, Alex;Sutskever, Ilya;Hinton, Geoffrey E.
通讯作者: Hinton, Geoffrey E.