Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel

Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Stanislav Fort;G. Dziugaite;Mansheej Paul;Sepideh Kharaghani;Daniel M. Roy;S. Ganguli
Stanislav Fort;G. Dziugaite;Mansheej Paul;Sepideh Kharaghani;Daniel M. Roy;S. Ganguli
中科院分区:
其他
文献类型:
--
作者:
Stanislav Fort;G. Dziugaite;Mansheej Paul;Sepideh Kharaghani;Daniel M. Roy;S. Ganguli

文献摘要

相似文献

在适当初始化的宽网络中,小学习率将深度神经网络(DNN)转变为神经正切核(NTK)机器,其训练动态可通过网络在初始化时的线性权重展开得到很好的近似。然而,标准训练与其线性化的偏离方式人们还知之甚少。我们研究非线性深度网络的训练动态、损失地貌的几何形状以及与数据相关的NTK的时间演化之间的关系。我们通过对训练进行大规模的现象学分析来实现这一点,综合了表征损失地貌几何形状和NTK动态的多种度量。在多种神经架构和数据集中,我们发现这些不同的度量以高度相关的方式演化,揭示了深度学习过程的一个通用图景。在这个图景中,深度网络训练呈现出一种高度混沌的快速初始瞬态,在2到3个轮次内就确定了包含训练终点的低损失的最终线性连接盆地。在这个混沌瞬态期间,NTK快速变化,从训练数据中学习有用的特征,使其能够在不到3到4个轮次内比标准的初始NTK性能高出3倍。在这个快速混沌瞬态之后,NTK以恒定速度变化,其性能在15%到45%的训练时间内与完整网络训练的性能相匹配。总体而言,我们的分析揭示了在训练时间上一组不同度量之间的显著相关性,这种相关性由最初几个轮次中从快速混沌到稳定的转变所支配,这一起为开发更准确的深度学习理论带来了挑战和机遇。
In suitably initialized wide networks, small learning rates transform deep neural networks (DNNs) into neural tangent kernel (NTK) machines, whose training dynamics is well-approximated by a linear weight expansion of the network at initialization. Standard training, however, diverges from its linearization in ways that are poorly understood. We study the relationship between the training dynamics of nonlinear deep networks, the geometry of the loss landscape, and the time evolution of a data-dependent NTK. We do so through a large-scale phenomenological analysis of training, synthesizing diverse measures characterizing loss landscape geometry and NTK dynamics. In multiple neural architectures and datasets, we find these diverse measures evolve in a highly correlated manner, revealing a universal picture of the deep learning process. In this picture, deep network training exhibits a highly chaotic rapid initial transient that within 2 to 3 epochs determines the final linearly connected basin of low loss containing the end point of training. During this chaotic transient, the NTK changes rapidly, learning useful features from the training data that enables it to outperform the standard initial NTK by a factor of 3 in less than 3 to 4 epochs. After this rapid chaotic transient, the NTK changes at constant velocity, and its performance matches that of full network training in 15% to 45% of training time. Overall, our analysis reveals a striking correlation between a diverse set of metrics over training time, governed by a rapid chaotic to stable transition in the first few epochs, that together poses challenges and opportunities for the development of more accurate theories of deep learning.