On the linearity of large non-linear models: when and why the tangent kernel is constant

On the linearity of large non-linear models: when and why the tangent kernel is constant
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Chaoyue Liu;Libin Zhu;M. Belkin
Chaoyue Liu;Libin Zhu;M. Belkin
中科院分区:
其他
文献类型:
--
作者:
Chaoyue Liu;Libin Zhu;M. Belkin

文献摘要

相似文献

这项工作的目标是揭示某些神经网络在其宽度接近无穷大时向线性过渡的显着现象。我们表明,过渡到线性的模型,等价地,恒定的(神经)切线内核(NTK)的结果从网络的海森矩阵的范数作为网络宽度的函数的标度特性。我们提出了一个通用的框架,通过适用于标准类的神经网络的海森尺度的切线核的恒定性。我们的分析提供了一个新的角度来看,常数正切核的现象,这是不同于广泛接受的“懒惰的训练”。此外,我们证明了向线性的过渡不是宽神经网络的一般性质,并且当网络的最后一层是非线性时不成立。它也不是通过梯度下降成功优化所必需的。
The goal of this work is to shed light on the remarkable phenomenon of transition to linearity of certain neural networks as their width approaches infinity. We show that the transition to linearity of the model and, equivalently, constancy of the (neural) tangent kernel (NTK) result from the scaling properties of the norm of the Hessian matrix of the network as a function of the network width. We present a general framework for understanding the constancy of the tangent kernel via Hessian scaling applicable to the standard classes of neural networks. Our analysis provides a new perspective on the phenomenon of constant tangent kernel, which is different from the widely accepted "lazy training". Furthermore, we show that the transition to linearity is not a general property of wide neural networks and does not hold when the last layer of the network is non-linear. It is also not necessary for successful optimization by gradient descent.