E-LSTM: An Efficient Hardware Architecture for Long Short-Term Memory
E-LSTM: An Efficient Hardware Architecture for Long Short-Term Memory
复制标题
E-LSTM:一种用于长短期记忆的高效硬件架构
DOI:
10.1109/jetcas.2019.2911739
复制
发表时间:
2019-04
影响因子:
4.6
通讯作者:
Wang Zhongfeng
中科院分区:
文献类型:
--
作者:
Wang Meiqi;Wang Zhisheng;Lu Jinming;Lin Jun;Wang Zhongfeng
Long Short-Term Memory (LSTM) and its variants have been widely adopted in many sequential learning tasks, such as speech recognition and machine translation. Significant accuracy improvements can be achieved using complex LSTM model with a large memory requirement and high computational complexity, which is time-consuming and energy demanding. The low-latency and energy-efficiency requirements of the real-world applications make model compression and hardware acceleration for LSTM an urgent need. In this paper, several hardware-efficient network compression schemes are introduced first, including structured top-<inline-formula> <tex-math notation="LaTeX">$k$ </tex-math></inline-formula> pruning, clipped gating, and multiplication-free quantization, to reduce the model size and the number of matrix operations by 32 <inline-formula> <tex-math notation="LaTeX">$\times $ </tex-math></inline-formula> and 21.6 <inline-formula> <tex-math notation="LaTeX">$\times $ </tex-math></inline-formula>, respectively, with negligible accuracy loss. Furthermore, efficient hardware architectures for accelerating the compressed LSTM are proposed, which support the inference of multi-layer and multiple time steps. The computation process is judiciously reorganized and the memory access pattern is well optimized, which alleviate the limited memory bandwidth bottleneck and enable higher throughput. Moreover, the parallel processing strategy is carefully designed to make full use of the sparsity introduced by pruning and clipped gating with high hardware utilization efficiency. Implemented on Intel Arria10 S<inline-formula> <tex-math notation="LaTeX">$\times $ </tex-math></inline-formula>660 FPGA running at 200MHz, the proposed design is able to achieve 1.4–2.2 <inline-formula> <tex-math notation="LaTeX">$\times $ </tex-math></inline-formula> energy efficiency and requires significantly less hardware resources compared with the state-of-the-art LSTM implementations.
登录
查看更多内容
影响因子:
3.9
作者:
Zhisheng Wang;Jun Lin;Zongfeng Wang
通讯作者:
Zongfeng Wang
DOI:
--
发表时间:
1993-02
期刊:
--
影响因子:
--
作者:
J. Garofolo;L. Lamel;W. Fisher;Jonathan G. Fiscus;D. S. Pallett;Nancy L. Dahlgren
通讯作者:
J. Garofolo;L. Lamel;W. Fisher;Jonathan G. Fiscus;D. S. Pallett;Nancy L. Dahlgren
DOI:
10.1145/3243176.3243184
发表时间:
2017-11
期刊:
Proceedings of the 27th International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
作者:
Franyell Silfa;Gem Dot;J. Arnau;Antonio González
通讯作者:
Franyell Silfa;Gem Dot;J. Arnau;Antonio González
DOI:
10.1109/tvlsi.2018.2819190
发表时间:
2018-07
影响因子:
2.8
作者:
Yun Long;Taesik Na;S. Mukhopadhyay
通讯作者:
Yun Long;Taesik Na;S. Mukhopadhyay
DOI:
--
发表时间:
2016-03
期刊:
ArXiv
影响因子:
--
作者:
D. Miyashita;Edward H. Lee;B. Murmann
通讯作者:
D. Miyashita;Edward H. Lee;B. Murmann