Theoretical Insights Into the Optimization Landscape of Over-Parameterized Shallow Neural Networks

Theoretical Insights Into the Optimization Landscape of Over-Parameterized Shallow Neural Networks
复制标题

DOI:
10.1109/tit.2018.2854560
复制
发表时间:
2019-02-01
影响因子:
2.5
通讯作者:
Lee, Jason D.
Lee, Jason D.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Soltanolkotabi, Mahdi;Javanmard, Adel;Lee, Jason D.

文献摘要

被引文献

相似文献

在本文中,我们研究了学习最适合培训数据集的浅层人工神经网络的问题。我们在过度参数化的制度中研究了这个问题,在该制度中,观测值的数量少于模型中的参数数量。我们表明,借助二次激活,训练的优化格局(这样的浅神经网络)具有某些有利的特征,可以使用各种局部搜索启发式学有效地找到全球最佳模型。该结果适用于输入/输出对的任意培训数据。对于可区分的激活功能,我们还表明,当适当初始化时,梯度下降以线性速率收敛到全球最佳模型。该结果着重于选择输入的可实现模型。根据高斯分布和标签是根据种植的重量系数产生的。
In this paper, we study the problem of learning a shallow artificial neural network that best fits a training data set. We study this problem in the over-parameterized regime where the numbers of observations are fewer than the number of parameters in the model. We show that with the quadratic activations, the optimization landscape of training, such shallow neural networks, has certain favorable characteristics that allow globally optimal models to be found efficiently using a variety of local search heuristics. This result holds for an arbitrary training data of input/output pairs. For differentiable activation functions, we also show that gradient descent, when suitably initialized, converges at a linear rate to a globally optimal model. This result focuses on a realizable model where the inputs are chosen i.i.d. from a Gaussian distribution and the labels are generated according to planted weight coefficients.