Scheduled Restart Momentum for Accelerated Stochastic Gradient Descent

Scheduled Restart Momentum for Accelerated Stochastic Gradient Descent
复制标题

DOI:
10.1137/21m1453311
复制
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Bao Wang;T. Nguyen;A. Bertozzi;Richard Baraniuk;S. Osher
Bao Wang;T. Nguyen;A. Bertozzi;Richard Baraniuk;S. Osher
中科院分区:
其他
文献类型:
--
作者:
Bao Wang;T. Nguyen;A. Bertozzi;Richard Baraniuk;S. Osher

文献摘要

被引文献

相似文献

具有恒定动量的随机梯度下降(SGD)及其变体(如Adam)是训练深度神经网络(DNN)的首选优化算法。由于DNN训练在计算上非常昂贵,因此人们对加速收敛非常感兴趣。Nesterov加速梯度(NAG)使用特殊设计的动量提高了凸优化的梯度下降(GD)的收敛速度;然而,当使用不精确的梯度时(如SGD),它会积累错误,最好的情况下会减慢收敛速度,最坏的情况下会发散。在本文中,我们提出了计划重启SGD(SRSGD),这是一种新的NAG风格的DNN训练方案。SRSGD用NAG中增加的动量代替SGD中的恒定动量,但根据时间表将动量重置为零来稳定迭代。使用各种图像分类模型和基准测试,我们证明,在训练DNN时,SRSGD显著提高了收敛性和泛化能力;例如,在训练ResNet 200进行ImageNet分类时,SRSGD的错误率为20.93%,而基准测试为22.13%。随着网络的深入,这些改进变得更加重要。此外,在CIFAR和ImageNet上,SRSGD达到了类似甚至更好的错误率,与SGD基线相比,训练时间明显减少。
Stochastic gradient descent (SGD) with constant momentum and its variants such as Adam are the optimization algorithms of choice for training deep neural networks (DNNs). Since DNN training is incredibly computationally expensive, there is great interest in speeding up the convergence. Nesterov accelerated gradient (NAG) improves the convergence rate of gradient descent (GD) for convex optimization using a specially designed momentum; however, it accumulates error when an inexact gradient is used (such as in SGD), slowing convergence at best and diverging at worst. In this paper, we propose Scheduled Restart SGD (SRSGD), a new NAG-style scheme for training DNNs. SRSGD replaces the constant momentum in SGD by the increasing momentum in NAG but stabilizes the iterations by resetting the momentum to zero according to a schedule. Using a variety of models and benchmarks for image classification, we demonstrate that, in training DNNs, SRSGD significantly improves convergence and generalization; for instance in training ResNet200 for ImageNet classification, SRSGD achieves an error rate of 20.93% vs. the benchmark of 22.13%. These improvements become more significant as the network grows deeper. Furthermore, on both CIFAR and ImageNet, SRSGD reaches similar or even better error rates with significantly fewer training epochs compared to the SGD baseline.