On the momentum term in gradient descent learning algorithms

On the momentum term in gradient descent learning algorithms
复制标题

DOI:
10.1016/s0893-6080(98)00116-6
复制
发表时间:
1999-01-01
期刊:
影响因子:
7.8
通讯作者:
Qian, N
Qian, N
中科院分区:
计算机科学1区
文献类型:
--
作者:
Qian, N

文献摘要

被引文献

相似文献

动量项通常包含在连接式学习算法的仿真中。虽然众所周知,这样的术语可以极大地提高学习速度,但对其机制的严谨研究却很少。在本文中,我证明了在连续时间的限制下,动量参数类似于牛顿粒子在守恒力场中通过粘性介质的质量。系统在局部极小值附近的行为等价于一组耦合的阻尼谐振子。动量项使系统的某些本征分量更接近临界阻尼,从而提高了收敛速度。对于计算机模拟中使用的离散时间情况,也可以得到类似的结果。特别地,我们得到了学习速率和动量参数的收敛的界,并证明了动量项可以增加系统收敛的学习速率的范围。并分析了收敛的最优条件。(C)1999爱思唯尔科学有限公司。保留所有权利。
A momentum term is usually included in the simulations of connectionist learning algorithms. Although it is well known that such a term greatly improves the speed of learning, there have been few rigorous studies of its: mechanisms. In this paper, I show that in the limit of continuous time, the momentum parameter is analogous to the mass of Newtonian particles that move through a viscous medium in a conservative force field. The behavior of the system near a local minimum is equivalent to a set of coupled and damped harmonic oscillators. The momentum term improves the speed of convergence by bringing some eigen components of the system closer to critical damping. Similar results can be obtained for the discrete time case used in computer simulations. In particular, I derive the bounds for convergence on learning-rate and momentum parameters, and demonstrate that the momentum term can increase the range of learning rate over which the system converges. The optimal condition for convergence is also analyzed. (C) 1999 Elsevier Science Ltd. All rights reserved.