What training reveals about neural network complexity

What training reveals about neural network complexity
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Andreas Loukas;Marinos Poiitis;S. Jegelka
Andreas Loukas;Marinos Poiitis;S. Jegelka
中科院分区:
其他
文献类型:
--
作者:
Andreas Loukas;Marinos Poiitis;S. Jegelka

文献摘要

被引文献

相似文献

这项工作探讨了良性训练假设(BTH),该假设认为深度神经网络(NN)正在学习的功能的复杂性可以通过其训练动态来推断。我们的分析提供了BTH的证据,通过将NN的Lipschitz常数在输入空间的不同区域与随机训练过程的行为相关联。我们首先观察到,接近训练数据的Lipschitz常数会影响参数轨迹的各个方面,更复杂的网络具有更长的轨迹、更大的方差,并且通常偏离初始化更远。然后,我们表明,第一层偏差训练得更稳定的NN(即,缓慢且变化很小)即使在远离任何训练点的输入空间的区域中也具有有限的复杂度。最后,我们发现,使用Dropout的稳定训练意味着一个依赖于训练和数据的泛化界,该泛化界随着参数的数量呈多项式增长。总的来说,我们的研究结果支持了这样一种直觉,即良好的训练行为可能是对良好泛化的有用偏见。
This work explores the Benevolent Training Hypothesis (BTH) which argues that the complexity of the function a deep neural network (NN) is learning can be deduced by its training dynamics. Our analysis provides evidence for BTH by relating the NN's Lipschitz constant at different regions of the input space with the behavior of the stochastic training procedure. We first observe that the Lipschitz constant close to the training data affects various aspects of the parameter trajectory, with more complex networks having a longer trajectory, bigger variance, and often veering further from their initialization. We then show that NNs whose 1st layer bias is trained more steadily (i.e., slowly and with little variation) have bounded complexity even in regions of the input space that are far from any training point. Finally, we find that steady training with Dropout implies a training- and data-dependent generalization bound that grows poly-logarithmically with the number of parameters. Overall, our results support the intuition that good training behavior can be a useful bias towards good generalization.