The general inefficiency of batch training for gradient descent learning

The general inefficiency of batch training for gradient descent learning
复制标题

DOI:
10.1016/s0893-6080(03)00138-2
复制
发表时间:
2003-12-01
期刊:
影响因子:
7.8
通讯作者:
Martinez, TR
Martinez, TR
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wilson, DR;Martinez, TR

文献摘要

被引文献

相似文献

神经网络的梯度下降训练可以以批处理或在线方式完成。在神经网络社区中,一个广为流传的神话是,批量训练比在线训练更快或更快和/或更“正确”,因为它使用了更好的真实梯度近似值来更新权重。本文解释了为什么批量训练几乎总是比在线训练慢,通常慢几个数量级,特别是在大型训练集上。主要原因是由于在线训练能够在每个时期跟踪误差曲面中的曲线,这使得它能够安全地使用更大的学习率,从而通过训练数据以更少的迭代收敛。大型(20,000个实例)语音识别任务和其他26个学习任务的实证结果表明,使用在线训练比批量训练可以更快地达到收敛,而且准确度没有明显差异。(C)2003爱思唯尔有限公司。保留所有权利。
Gradient descent training of neural networks can be done in either a batch or on-line manner. A widely held myth in the neural network community is that batch training is as fast or faster and/or more 'correct' than on-line training because it supposedly uses a better approximation of the true gradient for its weight updates. This paper explains why batch training is almost always slower than on-line training-often orders of magnitude slower---especially on large training sets. The main reason is due to the ability of on-line training to follow curves in the error surface throughout each epoch, which allows it to safely use a larger learning rate and thus converge with less iterations through the training data. Empirical results on a large (20,000-instance) speech recognition task and on 26 other learning tasks demonstrate that convergence can be reached significantly faster using on-line training than batch training, with no apparent difference in accuracy. (C) 2003 Elsevier Ltd. All rights reserved.