The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks

The Surprising Simplicity of the Early-Time Learning Dynamics of Neural Networks
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Wei Hu;Lechao Xiao;Ben Adlam;Jeffrey Pennington
Wei Hu;Lechao Xiao;Ben Adlam;Jeffrey Pennington
中科院分区:
其他
文献类型:
--
作者:
Wei Hu;Lechao Xiao;Ben Adlam;Jeffrey Pennington

文献摘要

被引文献

相似文献

现代神经网络通常被视为复杂的黑盒函数,由于其对数据的非线性依赖性和损失景观的非凸性,其行为难以理解。在这项工作中,我们表明,这些共同的看法可以在学习的早期阶段是完全错误的。特别是,我们正式证明,对于一类行为良好的输入分布,两层全连接神经网络的早期学习动态可以通过在输入上训练一个简单的线性模型来模仿。我们还认为,这种令人惊讶的简单性可以在具有更多层和卷积架构的网络中持续存在,我们根据经验验证了这一点。我们分析的关键是约束初始化时的神经切核(NTK)与数据核的仿射变换之间的差异的谱范数;然而,与使用NTK的许多先前结果不同,我们不要求网络具有不成比例的大宽度,并且允许网络在稍后的训练中逃离核机制。
Modern neural networks are often regarded as complex black-box functions whose behavior is difficult to understand owing to their nonlinear dependence on the data and the nonconvexity in their loss landscapes. In this work, we show that these common perceptions can be completely false in the early phase of learning. In particular, we formally prove that, for a class of well-behaved input distributions, the early-time learning dynamics of a two-layer fully-connected neural network can be mimicked by training a simple linear model on the inputs. We additionally argue that this surprising simplicity can persist in networks with more layers and with convolutional architecture, which we verify empirically. Key to our analysis is to bound the spectral norm of the difference between the Neural Tangent Kernel (NTK) at initialization and an affine transform of the data kernel; however, unlike many previous results utilizing the NTK, we do not require the network to have disproportionately large width, and the network is allowed to escape the kernel regime later in training.