Implicit Bias of Linear RNNs

Implicit Bias of Linear RNNs
复制标题

DOI:
--
复制
发表时间:
2021-01
期刊:
The Annals of Statistics
影响因子:
--
通讯作者:
M Motavali Emami;Mojtaba Sahraee-Ardakan;Parthe Pandit;S. Rangan;A. Fletcher
M Motavali Emami;Mojtaba Sahraee-Ardakan;Parthe Pandit;S. Rangan;A. Fletcher
中科院分区:
其他
文献类型:
--
作者:
M Motavali Emami;Mojtaba Sahraee-Ardakan;Parthe Pandit;S. Rangan;A. Fletcher

文献摘要

相似文献

基于经验研究的当代智慧表明,标准的递归神经网络(RNN)在需要长期记忆的任务中表现不佳。然而,对这一行为的准确推理仍不清楚。本文在线性RNN的特殊情况下对这一性质进行了严格的解释。虽然这项工作仅限于线性RNN,但即使是这些系统,由于其非线性的参数化,传统上也很难分析。利用最近发展的核制度分析,我们的主要结果表明,从随机初始化学习的线性RNN在功能上等价于某一加权一维卷积网络。重要的是,等效模型中的权重导致对卷积中具有较小时间滞后的元素的隐式偏差,从而导致较短的记忆。这种偏差的程度取决于初始化时转移核矩阵的方差,并且与经典的爆炸和消失梯度问题有关。该理论在合成数据和真实数据实验中都得到了验证。
Contemporary wisdom based on empirical studies suggests that standard recurrent neural networks (RNNs) do not perform well on tasks requiring long-term memory. However, precise reasoning for this behavior is still unknown. This paper provides a rigorous explanation of this property in the special case of linear RNNs. Although this work is limited to linear RNNs, even these systems have traditionally been difficult to analyze due to their non-linear parameterization. Using recently-developed kernel regime analysis, our main result shows that linear RNNs learned from random initializations are functionally equivalent to a certain weighted 1D-convolutional network. Importantly, the weightings in the equivalent model cause an implicit bias to elements with smaller time lags in the convolution and hence, shorter memory. The degree of this bias depends on the variance of the transition kernel matrix at initialization and is related to the classic exploding and vanishing gradients problem. The theory is validated in both synthetic and real data experiments.