Stochasticity of Deterministic Gradient Descent: Large Learning Rate for Multiscale Objective Function

Stochasticity of Deterministic Gradient Descent: Large Learning Rate for Multiscale Objective Function
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Lingkai Kong;Molei Tao
Lingkai Kong;Molei Tao
中科院分区:
其他
文献类型:
--
作者:
Lingkai Kong;Molei Tao

文献摘要

被引文献

相似文献

这篇文章表明,确定性梯度下降,不使用任何随机梯度近似,仍然可以表现出随机行为。特别是,它表明,如果目标函数表现出多尺度行为,那么在一个大的学习率制度,只解决宏观,但不是微观细节的目标,确定性GD动力学可以变得混乱和收敛到一个局部极小,但统计分布。一个充分条件也建立了近似这个长时间的统计限制的重新缩放吉布斯分布。提供理论和数值演示,理论部分依赖于使用有界噪声(而不是离散扩散)的随机映射的建设。
This article suggests that deterministic Gradient Descent, which does not use any stochastic gradient approximation, can still exhibit stochastic behaviors. In particular, it shows that if the objective function exhibit multiscale behaviors, then in a large learning rate regime which only resolves the macroscopic but not the microscopic details of the objective, the deterministic GD dynamics can become chaotic and convergent not to a local minimizer but to a statistical distribution. A sufficient condition is also established for approximating this long-time statistical limit by a rescaled Gibbs distribution. Both theoretical and numerical demonstrations are provided, and the theoretical part relies on the construction of a stochastic map that uses bounded noise (as opposed to discretized diffusions).