Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent

Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent
复制标题

DOI:
--
复制
发表时间:
2020-12
期刊:
--
影响因子:
--
通讯作者:
Kangqiao Liu;Liu Ziyin;Masakuni Ueda
Kangqiao Liu;Liu Ziyin;Masakuni Ueda
中科院分区:
其他
文献类型:
--
作者:
Kangqiao Liu;Liu Ziyin;Masakuni Ueda

文献摘要

相似文献

在学习率消失的情况下,随机梯度下降(SGD)现在已经得到了相对较好的理解。在这项工作中,我们建议研究SGD及其变体在非零学习率制度的基本属性。重点是推导出完全可解的结果,并讨论其影响。这项工作的主要贡献是推导出稳态分布离散时间SGD在二次损失函数有和没有动量,特别是,我们的结果的一个含义是,离散时间动态引起的波动采取扭曲的形状,是显着大于连续时间理论可以预测。在这项工作中考虑所提出的理论的应用程序的例子包括近似误差的变体SGD,小批量噪声的影响,最佳贝叶斯推理,逃逸率从一个尖锐的最小值,和平稳的协方差的一些二阶方法,包括阻尼牛顿的方法,自然梯度下降,亚当。
In the vanishing learning rate regime, stochastic gradient descent (SGD) is now relatively well understood. In this work, we propose to study the basic properties of SGD and its variants in the non-vanishing learning rate regime. The focus is on deriving exactly solvable results and discussing their implications. The main contributions of this work are to derive the stationary distribution for discrete-time SGD in a quadratic loss function with and without momentum; in particular, one implication of our result is that the fluctuation caused by discrete-time dynamics takes a distorted shape and is dramatically larger than a continuous-time theory could predict. Examples of applications of the proposed theory considered in this work include the approximation error of variants of SGD, the effect of minibatch noise, the optimal Bayesian inference, the escape rate from a sharp minimum, and the stationary covariance of a few second-order methods including damped Newton's method, natural gradient descent, and Adam.