LEARNING LONG-TERM DEPENDENCIES WITH GRADIENT DESCENT IS DIFFICULT

LEARNING LONG-TERM DEPENDENCIES WITH GRADIENT DESCENT IS DIFFICULT
复制标题

DOI:
10.1109/72.279181
复制
发表时间:
1994-03-01
影响因子:
--
通讯作者:
FRASCONI, P
FRASCONI, P
中科院分区:
其他
文献类型:
--
作者:
BENGIO, Y;SIMARD, P;FRASCONI, P

文献摘要

被引文献

相似文献

递归神经网络可以用于将输入序列映射到输出序列,例如用于识别,生产或预测问题。然而,在训练递归神经网络来执行任务时,已经报告了实际困难,其中输入/输出序列中存在的时间突发事件跨越长的间隔。我们展示了为什么基于梯度的学习算法面临着越来越困难的问题,因为要捕获的依赖关系的持续时间增加。这些结果揭示了梯度下降的有效学习和长时间锁定信息之间的权衡。基于对这个问题的理解,考虑了标准梯度下降的替代方案。
Recurrent neural networks can be used to map input sequences to output sequences, such as for recognition, production or prediction problems. However, practical difficulties have been reported in training recurrent neural networks to perform tasks in which the temporal contingencies present in the input/output sequences span long intervals. We show why gradient based learning algorithms face an increasingly difficult problem as the duration of the dependencies to be captured increases. These results expose a trade-off between efficient learning by gradient descent and latching on information for long periods. Based on an understanding of this problem, alternatives to standard gradient descent are considered.