The Multiscale Structure of Neural Network Loss Functions: The Effect on Optimization and Origin

The Multiscale Structure of Neural Network Loss Functions: The Effect on Optimization and Origin
复制标题

神经网络损失函数的多尺度结构:对优化和起源的影响

DOI:
10.48550/arxiv.2204.11326
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Lexing Ying
Lexing Ying
中科院分区:
--
文献类型:
--
作者:
Chao Ma;Lei Wu;Lexing Ying

文献摘要

参考文献

被引文献

相似文献

局部二次逼近已被广泛用于研究神经网络损失函数在最小值附近的优化。但是,它通常在最小值的一个非常小的邻域内成立,并且不能解释在优化过程中观察到的许多现象。在这项工作中,我们研究了神经网络损失函数的结构及其对优化的影响,在一个区域超出了良好的二次逼近。在数值上,我们观察到神经网络损失函数具有多尺度结构,表现在两个方面:(1)在最小值的邻域中,损失混合了连续的尺度并以次二次方式增长,(2)在更大的区域中,损失清楚地显示了几个独立的尺度。使用次二次增长,我们能够解释梯度下降(GD)方法观察到的稳定边缘现象[5]。利用分离尺度,我们通过简单的例子解释了学习率衰减的工作机制。最后,我们研究了多尺度结构的起源,并提出训练数据的非均匀性是其原因之一。通过构建一个两层神经网络问题,我们表明,不同幅度的训练数据引起不同规模的损失函数,产生次二次增长或多个单独的规模。
Local quadratic approximation has been extensively used to study the optimization of neural network loss functions around the minimum. Though, it usually holds in a very small neighborhood of the minimum, and cannot explain many phenomena observed during the optimization process. In this work, we study the structure of neural network loss functions and its implication on optimization in a region beyond the reach of good quadratic approximation. Numerically, we observe that neural network loss functions possesses a multiscale structure, manifested in two ways: (1) in a neighborhood of minima, the loss mixes a continuum of scales and grows subquadratically, and (2) in a larger region, the loss shows several separate scales clearly. Using the subquadratic growth, we are able to explain the Edge of Stability phenomenon [5] observed for gradient descent (GD) method. Using the separate scales, we explain the working mechanism of learning rate decay by simple examples. Finally, we study the origin of the multiscale structure and propose that the non-uniformity of training data is one of its cause. By constructing a two-layer neural network problem we show that training data with different magnitudes give rise to different scales of the loss function, producing subquadratic growth or multiple separate scales.
DOI: --
发表时间: 2020-02
期刊: ArXiv
影响因子: --
作者:
Lingkai Kong;Molei Tao
通讯作者: Lingkai Kong;Molei Tao
大学习率抑制同质性:收敛和平衡效应
DOI: --
发表时间: 2022
期刊: The International Conference on Learning Representations
影响因子: --
作者:
Wang, Yuqing;Chen, Minshuo;Zhao, Tuo;Tao, Molei
通讯作者: Tao, Molei
DOI: --
发表时间: 2022
期刊: PMLR
影响因子: --
作者:
Ahn, Kwangjun;Zhang, Jingzhao;Sra, Suvrit
通讯作者: Sra, Suvrit