Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers

Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
Albert Gu;Isys Johnson;Karan Goel;Khaled Kamal Saab;Tri Dao;A. Rudra;Christopher R'e
Albert Gu;Isys Johnson;Karan Goel;Khaled Kamal Saab;Tri Dao;A. Rudra;Christopher R'e
中科院分区:
其他
文献类型:
--
作者:
Albert Gu;Isys Johnson;Karan Goel;Khaled Kamal Saab;Tri Dao;A. Rudra;Christopher R'e

文献摘要

相似文献

递归神经网络(RNN)、时间卷积和神经微分方程(NDE)是一类流行的时间序列数据深度学习模型,它们在建模能力和计算效率方面都有独特的优势和折衷。我们介绍了一个受控制系统启发的简单序列模型,该模型概括了这些方法,同时解决了它们的缺点。线性状态空间层(LSSL)通过简单地模拟线性连续时间状态空间表示$\点{x}=Ax+Bu,y=Cx+Du$,将序列$u\mapstt映射到y$。理论上,我们证明了最小二乘模型与上述三类模型密切相关,并继承了它们的长处。例如,它们将卷积推广到连续时间,解释了常见的RNN启发式,并分享了NDE的特征,如时间尺度适应。然后,我们结合和推广了最近关于连续时间记忆的理论,引入了一个可训练的结构矩阵子集$A$,它赋予LSSL以长距离记忆。根据经验,将LSSL层堆叠到一个简单的深度神经网络中,可以在序列图像分类、真实医疗保健回归任务和语音中的长期相关性的时间序列基准测试中获得最先进的结果。在长度为16000序列的困难语音分类任务中,最小二乘算法比以前的方法高出24个准确点,甚至在100倍的短序列上优于使用手工特征的基线。
Recurrent neural networks (RNNs), temporal convolutions, and neural differential equations (NDEs) are popular families of deep learning models for time-series data, each with unique strengths and tradeoffs in modeling power and computational efficiency. We introduce a simple sequence model inspired by control systems that generalizes these approaches while addressing their shortcomings. The Linear State-Space Layer (LSSL) maps a sequence $u \mapsto y$ by simply simulating a linear continuous-time state-space representation $\dot{x} = Ax + Bu, y = Cx + Du$. Theoretically, we show that LSSL models are closely related to the three aforementioned families of models and inherit their strengths. For example, they generalize convolutions to continuous-time, explain common RNN heuristics, and share features of NDEs such as time-scale adaptation. We then incorporate and generalize recent theory on continuous-time memorization to introduce a trainable subset of structured matrices $A$ that endow LSSLs with long-range memory. Empirically, stacking LSSL layers into a simple deep neural network obtains state-of-the-art results across time series benchmarks for long dependencies in sequential image classification, real-world healthcare regression tasks, and speech. On a difficult speech classification task with length-16000 sequences, LSSL outperforms prior approaches by 24 accuracy points, and even outperforms baselines that use hand-crafted features on 100x shorter sequences.