Algorithm and Hardware Co-Design of Energy-Efficient LSTM Networks for Video Recognition With Hierarchical Tucker Tensor Decomposition

Algorithm and Hardware Co-Design of Energy-Efficient LSTM Networks for Video Recognition With Hierarchical Tucker Tensor Decomposition
复制标题

DOI:
10.1109/tc.2022.3212642
复制
发表时间:
2022-12
影响因子:
3.7
通讯作者:
Yu Gong;Miao Yin;Lingyi Huang;Chunhua Deng;Bo Yuan
Yu Gong;Miao Yin;Lingyi Huang;Chunhua Deng;Bo Yuan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yu Gong;Miao Yin;Lingyi Huang;Chunhua Deng;Bo Yuan

文献摘要

相似文献

长短期记忆(LSTM)是一种功能强大的深度神经网络,已广泛用于许多序列分析和建模应用。然而,LSTM网络的大模型问题使其实际部署仍然非常具有挑战性,特别是对于需要高维输入数据的视频识别任务。为了克服这一限制并充分释放LSTM模型的潜力,本文提出对高性能节能LSTM网络进行算法和硬件协同设计。在算法层面,我们建议开发基于完全分解分层塔克(FDHT)结构的LSTM,即FDHT-LSTM,它具有超低的模型复杂度,同时仍然具有高精度。为了充分获得这种有吸引力的算法优势,我们进一步开发了相应的定制硬件架构,以支持所提出的FDHT-LSTM模型的有效执行。通过对存储器访问机制的精心设计,底层硬件可以在不发生访问冲突的情况下,有效地支持复杂的矩阵变换。我们的评估结果表明,所提出的超紧凑FDHT-LSTM模型和相应的硬件加速器都实现了非常高的性能。与最先进的压缩LSTM模型相比,FDHT-LSTM在不同的视频识别数据集上既可以减少模型大小的数量级(超过1000 × 1000),又可以显著提高准确率(0.6%至12.7%)。同时,与最先进的张量分解面向模型的硬件TIE相比,我们提出的FDHT-LSTM架构在LSTM-Youtube工作负载上的吞吐量,面积效率和能源效率分别提高了2.5\times$2.5 ×,1.46\times$1.46 ×和2.41\times$2.41 ×。对于LSTM-UCF工作负载,我们提出的设计也优于TIE,吞吐量高出1.9\times1.9倍,能效高出1.83\times1.83倍,面积效率相当。
Long short-term memory (LSTM) is a type of powerful deep neural network that has been widely used in many sequence analysis and modeling applications. However, the large model size problem of LSTM networks make their practical deployment still very challenging, especially for the video recognition tasks that require high-dimensional input data. Aiming to overcome this limitation and fully unlock the potentials of LSTM models, in this paper we propose to perform algorithm and hardware co-design towards high-performance energy-efficient LSTM networks. At algorithm level, we propose to develop fully decomposed hierarchical Tucker (FDHT) structure-based LSTM, namely FDHT-LSTM, which enjoys ultra-low model complexity while still achieving high accuracy. In order to fully reap such attractive algorithmic benefit, we further develop the corresponding customized hardware architecture to support the efficient execution of the proposed FDHT-LSTM model. With the delicate design of memory access scheme, the complicated matrix transformation can be efficiently supported by the underlying hardware without any access conflict in an on-the-fly way. Our evaluation results show that both the proposed ultra-compact FDHT-LSTM models and the corresponding hardware accelerator achieve very high performance. Compared with the state-of-the-art compressed LSTM models, FDHT-LSTM enjoys both order-of-magnitude reduction (more than $1000 \times$1000×) in model size and significant accuracy improvement (0.6% to 12.7%) across different video recognition datasets. Meanwhile, compared with the state-of-the-art tensor decomposed model-oriented hardware TIE, our proposed FDHT-LSTM architecture achieve $2.5\times$2.5×, $1.46\times$1.46× and $2.41\times$2.41× increase in throughput, area efficiency and energy efficiency, respectively on LSTM-Youtube workload. For LSTM-UCF workload, our proposed design also outperforms TIE with $1.9\times$1.9× higher throughput, $1.83\times$1.83× higher energy efficiency and comparable area efficiency.