A Local Convergence Theory for Mildly Over-Parameterized Two-Layer Neural Network

A Local Convergence Theory for Mildly Over-Parameterized Two-Layer Neural Network
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
--
影响因子:
--
通讯作者:
Mo Zhou;Rong Ge;Chi Jin
Mo Zhou;Rong Ge;Chi Jin
中科院分区:
其他
文献类型:
--
作者:
Mo Zhou;Rong Ge;Chi Jin

文献摘要

被引文献

相似文献

虽然过度参数化被广泛认为是神经网络优化成功的关键,但大多数现有的过度参数化理论都没有完全解释原因-它们要么在神经切线核机制中工作,其中神经元不会移动太多,要么需要大量的神经元。在实践中,当使用教师神经网络生成数据时,即使是轻度过参数化的神经网络也可以实现0损失并恢复教师神经元的方向。本文对轻度过参数双层神经网络建立了一个局部收敛理论。我们证明,只要损失已经低于阈值(相关参数的多项式),过参数化的两层神经网络中的所有学生神经元将收敛到一个教师神经元,损失将变为0。我们的结果适用于任何数量的学生神经元,只要它至少与教师神经元的数量一样大,并且我们的收敛速度与学生神经元的数量无关。我们分析的一个关键组成部分是局部优化景观的新特征-我们表明梯度满足Lojasiewicz性质的特殊情况,这与以前工作中使用的局部强凸性或PL条件不同。
While over-parameterization is widely believed to be crucial for the success of optimization for the neural networks, most existing theories on over-parameterization do not fully explain the reason -- they either work in the Neural Tangent Kernel regime where neurons don't move much, or require an enormous number of neurons. In practice, when the data is generated using a teacher neural network, even mildly over-parameterized neural networks can achieve 0 loss and recover the directions of teacher neurons. In this paper we develop a local convergence theory for mildly over-parameterized two-layer neural net. We show that as long as the loss is already lower than a threshold (polynomial in relevant parameters), all student neurons in an over-parameterized two-layer neural network will converge to one of teacher neurons, and the loss will go to 0. Our result holds for any number of student neurons as long as it is at least as large as the number of teacher neurons, and our convergence rate is independent of the number of student neurons. A key component of our analysis is the new characterization of local optimization landscape -- we show the gradient satisfies a special case of Lojasiewicz property which is different from local strong convexity or PL conditions used in previous work.