adaQN: An Adaptive Quasi-Newton Algorithm for Training RNNs

adaQN: An Adaptive Quasi-Newton Algorithm for Training RNNs
复制标题

adaQN:用于训练 RNN 的自适应拟牛顿算法

DOI:
10.1007/978-3-319-46128-1_1
复制
发表时间:
2015
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Berahas
A. Berahas
中科院分区:
--
文献类型:
--
作者:
N. Keskar;A. Berahas

文献摘要

被引文献

相似文献

循环神经网络 (RNN) 是功能强大的模型,可以在多种模式识别问题上实现卓越的性能。然而,由于众所周知的“消失/爆炸”梯度问题,RNN 的训练在计算上是一项困难的任务。提出的用于训练 RNN 的算法要么不利用(或有限的)曲率信息并且具有较低的每次迭代复杂度,要么尝试以增加每次迭代成本为代价来获得重要的曲率信息。前一组包括对角缩放的一阶方法,例如 ADAGRAD 和 ADAM,而后者包括二阶算法,例如 Hessian-Free Newton 和 K-FAC。在本文中,我们提出了 adaQN,一种用于训练 RNN 的随机拟牛顿算法。我们的方法保留了较低的每次迭代成本,同时允许通过随机 L-BFGS 更新方案进行非对角缩放。该方法使用新颖的 L-BFGS 缩放初始化方案,并且明智地存储和保留 L-BFGS 曲率对。我们对两种语言建模任务进行了数值实验,并表明 adaQN 与流行的 RNN 训练算法具有竞争力。
Recurrent Neural Networks (RNNs) are powerful models that achieve exceptional performance on several pattern recognition problems. However, the training of RNNs is a computationally difficult task owing to the well-known "vanishing/exploding" gradient problem. Algorithms proposed for training RNNs either exploit no (or limited) curvature information and have cheap per-iteration complexity, or attempt to gain significant curvature information at the cost of increased per-iteration cost. The former set includes diagonally-scaled first-order methods such as ADAGRAD and ADAM, while the latter consists of second-order algorithms like Hessian-Free Newton and K-FAC. In this paper, we present adaQN, a stochastic quasi-Newton algorithm for training RNNs. Our approach retains a low per-iteration cost while allowing for non-diagonal scaling through a stochastic L-BFGS updating scheme. The method uses a novel L-BFGS scaling initialization scheme and is judicious in storing and retaining L-BFGS curvature pairs. We present numerical experiments on two language modeling tasks and show that adaQN is competitive with popular RNN training algorithms.