AEGD: adaptive gradient descent with energy

AEGD: adaptive gradient descent with energy
复制标题

DOI:
10.3934/naco.2023015
复制
发表时间:
2020-10
期刊:
Numerical Algebra, Control and Optimization
影响因子:
--
通讯作者:
Hailiang Liu;Xuping Tian
Hailiang Liu;Xuping Tian
中科院分区:
其他
文献类型:
--
作者:
Hailiang Liu;Xuping Tian

文献摘要

相似文献

我们提出了AEGD,一个新的算法的一阶梯度为基础的优化非凸目标函数,基于一个动态更新的能量变量。该方法被证明是无条件的能量稳定,无论步长大小。我们证明了非凸和凸目标的AEGD的能量依赖的收敛速度,这对于一个适当的小步长恢复所需的收敛速度的批量梯度下降。我们还提供了一个能量依赖的边界上的平稳收敛的AEGD的随机非凸设置。该方法易于实现,并且几乎不需要调整超参数。实验结果表明,AEGD适用于各种各样的优化问题:它是强大的初始数据,能够快速初始化的进展。对于深度神经网络,随机AEGD显示出与具有动量的SGD相当且通常更好的泛化性能。
We propose AEGD, a new algorithm for first-order gradient-based optimization of non-convex objective functions, based on a dynamically updated energy variable. The method is shown to be unconditionally energy stable, irrespective of the step size. We prove energy-dependent convergence rates of AEGD for both non-convex and convex objectives, which for a suitably small step size recovers desired convergence rates for the batch gradient descent. We also provide an energy-dependent bound on the stationary convergence of AEGD in the stochastic non-convex setting. The method is straightforward to implement and requires little tuning of hyper-parameters. Experimental results demonstrate that AEGD works well for a large variety of optimization problems: it is robust with respect to initial data, capable of making rapid initial progress. The stochastic AEGD shows comparable and often better generalization performance than SGD with momentum for deep neural networks.