Understanding Gradient Descent on Edge of Stability in Deep Learning

Understanding Gradient Descent on Edge of Stability in Deep Learning
复制标题

理解深度学习中稳定性边缘的梯度下降

DOI:
--
复制
发表时间:
2022
期刊:
arXiv.org
影响因子:
--
通讯作者:
A. Panigrahi
A. Panigrahi
中科院分区:
--
文献类型:
--
作者:
Sanjeev Arora;Zhiyuan Li;A. Panigrahi

文献摘要

被引文献

相似文献

Cohen等人的深度学习实验。[2021]当学习率(LR)和清晰度(即Hessian的最大特征值)不再像传统优化中那样表现时,确定性梯度下降(GD)显示出稳定的边缘(Eos)阶段。清晰度稳定在2美元/$LR左右,损失在迭代中上下波动,但总体趋势仍是下降的。本文从数学上分析了EOS阶段隐式正则化的一种新机制,即在最小损失流形上,由于非光滑损失而导致的GD更新在某种确定性流动中演化。这与之前许多关于隐式偏差的结果形成了鲜明对比,这些结果要么依赖于无限小的更新,要么依赖于渐变中的噪声。形式上,对于任何具有一定正则性条件的光滑函数$L$,对于(1)归一化的GD,即具有变化的LR$\Eta_t=FRAC{\Eta_t=FRAC{\ETa}{x(T)}的GD,损失$L$;(2)具有恒定的LR和损失$\Sqrt{L-\min_x L(X)}$的GD。两者都可以证明进入稳定的边缘,流形上的相关流使$\lambda_{1}(\nabla^2 L)$最小化。上述理论结果得到了实验研究的证实。
Deep learning experiments by Cohen et al. [2021] using deterministic Gradient Descent (GD) revealed an Edge of Stability (EoS) phase when learning rate (LR) and sharpness (i.e., the largest eigenvalue of Hessian) no longer behave as in traditional optimization. Sharpness stabilizes around $2/$LR and loss goes up and down across iterations, yet still with an overall downward trend. The current paper mathematically analyzes a new mechanism of implicit regularization in the EoS phase, whereby GD updates due to non-smooth loss landscape turn out to evolve along some deterministic flow on the manifold of minimum loss. This is in contrast to many previous results about implicit bias either relying on infinitesimal updates or noise in gradient. Formally, for any smooth function $L$ with certain regularity condition, this effect is demonstrated for (1) Normalized GD, i.e., GD with a varying LR $\eta_t =\frac{\eta}{\| \nabla L(x(t)) \|}$ and loss $L$; (2) GD with constant LR and loss $\sqrt{L- \min_x L(x)}$. Both provably enter the Edge of Stability, with the associated flow on the manifold minimizing $\lambda_{1}(\nabla^2 L)$. The above theoretical results have been corroborated by an experimental study.