Dynamics of Deep Neural Networks and Neural Tangent Hierarchy

Dynamics of Deep Neural Networks and Neural Tangent Hierarchy
复制标题

DOI:
--
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Jiaoyang Huang;H. Yau
Jiaoyang Huang;H. Yau
中科院分区:
其他
文献类型:
--
作者:
Jiaoyang Huang;H. Yau

文献摘要

被引文献

相似文献

通过梯度下降训练的深度神经网络的演化可以用[20]中介绍的神经切线核(NTK)来描述,在[20]中证明了在无限宽度极限下,NTK收敛于显式极限核,并且在训练过程中保持不变。最近的一些论文也隐含了NTK[6,13,14]。在过度参数化状态下,一个完全训练的深度神经网络确实等同于使用极限NTK的核回归预测器。梯度下降法使深度过参数化神经网络的训练损失为零。然而,在[5]中观察到,使用限制NTK的核回归与深度神经网络之间存在性能差距。这种性能差距很可能是由于有限宽度效应导致的训练过程中NTK的变化。NTK在训练过程中的变化是描述深度神经网络泛化特征的核心。在本文中,我们研究了有限宽度深度全连接神经网络的NTK动态。我们推导了一个无限层次的常微分方程,即神经切线层次(NTH),它捕捉了深度神经网络的梯度下降动态。此外,在一定的神经网络宽度和数据集维数条件下,我们证明了NTH的截断层次可以在任意精度上近似NTK的动态。这种描述使得直接研究深度神经网络的NTK变化成为可能,并阐明了使用相应的限制NTK,深度神经网络优于核回归的观察。
The evolution of a deep neural network trained by the gradient descent can be described by its neural tangent kernel (NTK) as introduced in [20], where it was proven that in the infinite width limit the NTK converges to an explicit limiting kernel and it stays constant during training. The NTK was also implicit in some other recent papers [6,13,14]. In the overparametrization regime, a fully-trained deep neural network is indeed equivalent to the kernel regression predictor using the limiting NTK. And the gradient descent achieves zero training loss for a deep overparameterized neural network. However, it was observed in [5] that there is a performance gap between the kernel regression using the limiting NTK and the deep neural networks. This performance gap is likely to originate from the change of the NTK along training due to the finite width effect. The change of the NTK along the training is central to describe the generalization features of deep neural networks. In the current paper, we study the dynamic of the NTK for finite width deep fully-connected neural networks. We derive an infinite hierarchy of ordinary differential equations, the neural tangent hierarchy (NTH) which captures the gradient descent dynamic of the deep neural network. Moreover, under certain conditions on the neural network width and the data set dimension, we prove that the truncated hierarchy of NTH approximates the dynamic of the NTK up to arbitrary precision. This description makes it possible to directly study the change of the NTK for deep neural networks, and sheds light on the observation that deep neural networks outperform kernel regressions using the corresponding limiting NTK.