Long short-term memory

Long short-term memory
复制标题

DOI:
10.1007/978-3-642-24797-2
复制
发表时间:
1997-11-15
期刊:
影响因子:
2.9
通讯作者:
Schmidhuber, J
Schmidhuber, J
中科院分区:
计算机科学4区
文献类型:
--
作者:
Hochreiter, S;Schmidhuber, J

文献摘要

被引文献

相似文献

学习通过循环反向传播在延长的时间间隔内存储信息需要很长时间,主要是因为不充分,衰减的错误回流。我们简要回顾了Hochreiter(1991)对这个问题的分析,然后通过引入一种称为长短期记忆(LSTM)的新颖,高效,基于梯度的方法来解决这个问题。在不造成伤害的情况下截断梯度,LSTM可以学习通过在特殊单元内的恒定误差旋转木马强制执行恒定误差,来桥接超过1000个离散时间步骤的最小时滞。乘法门单元学习打开和关闭,获得恒定的误差流。LSTM在空间和时间上都是局部的;其每个时间步和权重的计算复杂度为O(1)。我们的人工数据的实验涉及本地,分布式,实值和嘈杂的模式表示。与实时递归学习、时间反向传播、递归级联相关、Elman网络和神经序列分块相比,LSTM可以更成功地运行,并且学习速度更快。LSTM还可以解决复杂的人工长时间滞后任务,这些任务以前的递归网络算法从未解决过。
Learning to store information over extended time intervals by recurrent backpropagation takes a very long time, mostly because of insufficient, decaying error backflow. We briefly review Hochreiter's (1991) analysis of this problem, then address it by introducing a novel, efficient, gradient-based method called long short-term memory (LSTM). Truncating the gradient where this does not do harm, LSTM can learn to bridge minimal time lags in excess of 1000 discrete-time steps by enforcing constant error now through constant error carousels within special units. Multiplicative gate units learn to open and close access to the constant error flow. LSTM is local in space and time; its computational complexity per time step and weight is O(1). Our experiments with artificial data involve local, distributed, real-valued, and noisy pattern representations. In comparisons with real-time recurrent learning, back propagation through time, recurrent cascade correlation, Elman nets, and neural sequence chunking, LSTM leads to many more successful runs, and learns much faster. LSTM also solves complex, artificial long-time-lag tasks that have never been solved by previous recurrent network algorithms.