Convolutional Tensor-Train LSTM for Spatio-temporal Learning

Convolutional Tensor-Train LSTM for Spatio-temporal Learning
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Jiahao Su;Wonmin Byeon;Furong Huang;J. Kautz;Anima Anandkumar
Jiahao Su;Wonmin Byeon;Furong Huang;J. Kautz;Anima Anandkumar
中科院分区:
其他
文献类型:
--
作者:
Jiahao Su;Wonmin Byeon;Furong Huang;J. Kautz;Anima Anandkumar

文献摘要

被引文献

相似文献

从时空数据中学习有许多应用,如人类行为分析、对象跟踪、视频压缩和物理模拟。然而,现有的方法在具有挑战性的视频任务(如长期预测)中仍然表现不佳。这是因为这类具有挑战性的任务需要学习视频序列中的长期时空相关性。在本文中,我们提出了一个高阶卷积LSTM模型,可以有效地学习这些相关性,以及历史的简洁表示。这是通过一个新颖的张量训练模块来完成的,该模块通过组合卷积特征来进行预测。为了使这在计算和内存需求方面可行,我们提出了一种新的高阶模型的卷积张量序列分解。这种分解通过联合逼近卷积核序列作为低秩张量-训练分解来降低模型复杂性。因此,我们的模型优于现有的方法,但只使用了一小部分参数,包括基线模型。我们的结果在广泛的应用和数据集中实现了最先进的性能,包括在moving - mist -2和KTH动作数据集上的多步视频预测,以及在Something-Something V2数据集上的早期活动识别。
Learning from spatio-temporal data has numerous applications such as human-behavior analysis, object tracking, video compression, and physics simulation.However, existing methods still perform poorly on challenging video tasks such as long-term forecasting. This is because these kinds of challenging tasks require learning long-term spatio-temporal correlations in the video sequence. In this paper, we propose a higher-order convolutional LSTM model that can efficiently learn these correlations, along with a succinct representations of the history. This is accomplished through a novel tensor train module that performs prediction by combining convolutional features across time. To make this feasible in terms of computation and memory requirements, we propose a novel convolutional tensor-train decomposition of the higher-order model. This decomposition reduces the model complexity by jointly approximating a sequence of convolutional kernels asa low-rank tensor-train factorization. As a result, our model outperforms existing approaches, but uses only a fraction of parameters, including the baseline models.Our results achieve state-of-the-art performance in a wide range of applications and datasets, including the multi-steps video prediction on the Moving-MNIST-2and KTH action datasets as well as early activity recognition on the Something-Something V2 dataset.