Towards Extremely Compact RNNs for Video Recognition with Fully Decomposed Hierarchical Tucker Structure

Towards Extremely Compact RNNs for Video Recognition with Fully Decomposed Hierarchical Tucker Structure
复制标题

DOI:
10.1109/cvpr46437.2021.01191
复制
发表时间:
2021-04
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Miao Yin;Siyu Liao;Xiao-Yang Liu;Xiaodong Wang;Bo Yuan
Miao Yin;Siyu Liao;Xiao-Yang Liu;Xiaodong Wang;Bo Yuan
中科院分区:
其他
文献类型:
--
作者:
Miao Yin;Siyu Liao;Xiao-Yang Liu;Xiaodong Wang;Bo Yuan

文献摘要

被引文献

相似文献

递归神经网络(RNN)在序列分析和建模中有着广泛的应用。然而,在处理高维数据时,RNN通常需要非常大的模型大小,从而带来一系列部署挑战。虽然已经提出了各种减少RNN模型规模的工作,但在资源受限的环境中执行RNN模型仍然是一个非常具有挑战性的问题。在本文中,我们建议建立具有完全分解分层Tucker(FDHT)结构的极其紧凑的RNN模型。与其他张量分解方法相比,HT分解不仅具有更高的存储代价,而且对紧致RNN模型的精度性能有更好的改善。同时,与现有的基于张量分解的方法只能分解RNN的输入到隐含层不同,我们提出的完全分解方法能够在保持很高精度的情况下对整个RNN模型进行全面的压缩。我们在几个流行的视频识别数据集上的实验结果表明,我们提出的基于完全分解的层次Tucker的LSTM(FDHT-LSTM)是非常紧凑和高效的。据我们所知,FDHT-LSTM第一次在不同的数据集上仅用几千个参数(3,132到8,808个)就能始终如一地达到非常高的精度。与目前最先进的压缩RNN模型如TT-LSTM、TR-LSTM和BT-LSTM相比,我们的FDHT-LSTM同时具有更少的参数数量级(3985×到10,711×)和显著的精度提高(0.6%到12.7%)。
Recurrent Neural Networks (RNNs) have been widely used in sequence analysis and modeling. However, when processing high-dimensional data, RNNs typically require very large model sizes, thereby bringing a series of deployment challenges. Although various prior works have been proposed to reduce the RNN model sizes, executing RNN models in the resource-restricted environments is still a very challenging problem. In this paper, we propose to develop extremely compact RNN models with fully decomposed hierarchical Tucker (FDHT) structure. The HT decomposition does not only provide much higher storage cost reduction than the other tensor decomposition approaches, but also brings better accuracy performance improvement for the compact RNN models. Meanwhile, unlike the existing tensor decomposition-based methods that can only decompose the input-to-hidden layer of RNNs, our proposed fully decomposition approach enables the comprehensive compression for the entire RNN models with maintaining very high accuracy. Our experimental results on several popular video recognition datasets show that, our proposed fully decomposed hierarchical tucker-based LSTM (FDHT-LSTM) is extremely compact and highly efficient. To the best of our knowledge, FDHT-LSTM, for the first time, consistently achieves very high accuracy with only few thousand parameters (3,132 to 8,808) on different datasets. Compared with the state-of-the-art compressed RNN models, such as TT-LSTM, TR-LSTM and BT-LSTM, our FDHT-LSTM simultaneously enjoys both order-of-magnitude (3,985× to 10,711×) fewer parameters and significant accuracy improvement (0.6% to 12.7%).