Adaptive momentum with discriminative weight for neural network stochastic optimization

Adaptive momentum with discriminative weight for neural network stochastic optimization
复制标题

DOI:
10.1002/int.22854
复制
发表时间:
2022-09
影响因子:
7
通讯作者:
Jiyang Bai;Yuxiang Ren;Jiawei Zhang
Jiyang Bai;Yuxiang Ren;Jiawei Zhang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jiyang Bai;Yuxiang Ren;Jiawei Zhang

文献摘要

相似文献

具有动量的优化算法由于其快速收敛速度而被广泛用于构建深度学习模型。动量有助于加速参数更新过程中相关方向上的随机梯度下降,减小参数更新路径的振荡。动量优化算法中每一步的梯度都是由一部分训练样本计算的,存在随机性,可能给参数更新带来误差。在这种情况下,以固定权重将上一步的影响置于当前步的动量显然是不准确的,这传播了误差并阻碍了当前步的校正。此外,这样的超参数在应用中也很难调整。本文介绍了一种新的优化算法,即自适应动量判别加权法(DEAM)。DEAM提出了基于判别角自动计算动量项权重的方法,而不是用固定的超参数来分配动量项权重。动量项权重将被分配一个适当的值,用于配置当前步骤中的动量。这样,DEAM涉及更少的超参数。DEAM还包含一个新的回溯项,当需要校正最后一步时,该项限制了冗余更新。回溯项可以有效地自适应学习速率,并实现预期的更新。大量实验表明,DEAM在训练凸和非凸情况下的深度学习模型时,可以实现比现有优化算法更快的收敛速度。
Optimization algorithms with momentum have been widely used for building deep learning models because of the fast convergence rate. Momentum helps accelerate Stochastic gradient descent in relevant directions in parameter updating, minifying the oscillations of the parameters update route. The gradient of each step in optimization algorithms with momentum is calculated by a part of the training samples, so there exists stochasticity, which may bring errors to parameter updates. In this case, momentum placing the influence of the last step to the current step with a fixed weight is obviously inaccurate, which propagates the error and hinders the correction of the current step. Besides, such a hyperparameter can be extremely hard to tune in applications as well. In this paper, we introduce a novel optimization algorithm, namely, Discriminative wEight on Adaptive Momentum (DEAM). Instead of assigning the momentum term weight with a fixed hyperparameter, DEAM proposes to compute the momentum weight automatically based on the discriminative angle. The momentum term weight will be assigned with an appropriate value that configures momentum in the current step. In this way, DEAM involves fewer hyperparameters. DEAM also contains a novel backtrack term, which restricts redundant updates when the correction of the last step is needed. The backtrack term can effectively adapt the learning rate and achieve the anticipatory update as well. Extensive experiments demonstrate that DEAM can achieve a faster convergence rate than the existing optimization algorithms in training the deep learning models of both convex and nonconvex situations.