Eigenvalue Normalized Recurrent Neural Networks for Short Term Memory

Eigenvalue Normalized Recurrent Neural Networks for Short Term Memory
复制标题

DOI:
10.1609/aaai.v34i04.5831
复制
发表时间:
2019-11
期刊:
--
影响因子:
--
通讯作者:
Kyle E. Helfrich;Q. Ye
Kyle E. Helfrich;Q. Ye
中科院分区:
其他
文献类型:
--
作者:
Kyle E. Helfrich;Q. Ye

文献摘要

相似文献

最近开发了具有正交或酉递归矩阵的递归神经网络(RNN)的几种变体,以减轻消失/爆炸梯度问题并对序列的长期依赖性进行建模。然而,由于递归矩阵的特征值在单位圆上,递归状态保留了所有的输入信息,这可能不必要地消耗模型容量。在本文中,我们通过提出一种架构来解决这个问题,该架构扩展了正交/酉RNN,其状态由具有单位圆盘中特征值的递归矩阵生成。任何输入都会随着时间的推移而消散,并被新的输入所取代,模拟短期记忆。一个梯度下降算法推导出学习这样的递归矩阵。由此产生的方法,称为特征值归一化RNN(ENRNN),在几个实验中显示出高度的竞争力。
Several variants of recurrent neural networks (RNNs) with orthogonal or unitary recurrent matrices have recently been developed to mitigate the vanishing/exploding gradient problem and to model long-term dependencies of sequences. However, with the eigenvalues of the recurrent matrix on the unit circle, the recurrent state retains all input information which may unnecessarily consume model capacity. In this paper, we address this issue by proposing an architecture that expands upon an orthogonal/unitary RNN with a state that is generated by a recurrent matrix with eigenvalues in the unit disc. Any input to this state dissipates in time and is replaced with new inputs, simulating short-term memory. A gradient descent algorithm is derived for learning such a recurrent matrix. The resulting method, called the Eigenvalue Normalized RNN (ENRNN), is shown to be highly competitive in several experiments.