E-LSTM: An Efficient Hardware Architecture for Long Short-Term Memory

E-LSTM: An Efficient Hardware Architecture for Long Short-Term Memory
复制标题

E-LSTM:一种用于长短期记忆的高效硬件架构

DOI:
10.1109/jetcas.2019.2911739
复制
发表时间:
2019-04
影响因子:
4.6
通讯作者:
Wang Zhongfeng
Wang Zhongfeng
中科院分区:
工程技术2区
文献类型:
--
作者:
Wang Meiqi;Wang Zhisheng;Lu Jinming;Lin Jun;Wang Zhongfeng

文献摘要

参考文献

被引文献

相似文献

长短期记忆(LSTM)及其变体已被广泛应用于许多顺序学习任务,如语音识别和机器翻译。使用复杂的LSTM模型可以实现显著的精度提高,具有大的内存需求和高计算复杂度,这是耗时和耗能的要求。现实世界应用的低延迟和能效要求使得LSTM的模型压缩和硬件加速成为迫切需要。在本文中,首先介绍了几种硬件高效的网络压缩方案,包括结构化的top-<inline-formula><tex-math notation="LaTeX">$k$</tex-math></inline-formula>剪枝,裁剪门控,和乘法免费量化,以减少模型的大小和矩阵运算的数量分别为32<inline-formula><tex-math notation="LaTeX">$\times $</tex-math></inline-formula>和21.6<inline-formula><tex-math notation="LaTeX">$\times $</tex-math></inline-formula>,可以忽略不计的精度损失。此外,提出了用于加速压缩LSTM的有效硬件架构,该架构支持多层和多时间步的推理。该算法对计算过程进行了合理的重组,并对内存访问模式进行了优化,从而缓解了有限的内存带宽瓶颈,提高了吞吐量。同时,设计了并行处理策略,充分利用了剪枝和限幅门带来的稀疏性,提高了硬件利用率。在运行于200 MHz的Intel Arria 10 S<inline-formula><tex-math notation="LaTeX">$\times $660</tex-math></inline-formula> FPGA上实现,所提出的设计能够实现1.4-2.2<inline-formula><tex-math notation="LaTeX">$\times $的</tex-math></inline-formula>能效,并且与最先进的LSTM实现相比,所需的硬件资源显著减少。
Long Short-Term Memory (LSTM) and its variants have been widely adopted in many sequential learning tasks, such as speech recognition and machine translation. Significant accuracy improvements can be achieved using complex LSTM model with a large memory requirement and high computational complexity, which is time-consuming and energy demanding. The low-latency and energy-efficiency requirements of the real-world applications make model compression and hardware acceleration for LSTM an urgent need. In this paper, several hardware-efficient network compression schemes are introduced first, including structured top-<inline-formula> <tex-math notation="LaTeX">$k$ </tex-math></inline-formula> pruning, clipped gating, and multiplication-free quantization, to reduce the model size and the number of matrix operations by 32 <inline-formula> <tex-math notation="LaTeX">$\times $ </tex-math></inline-formula> and 21.6 <inline-formula> <tex-math notation="LaTeX">$\times $ </tex-math></inline-formula>, respectively, with negligible accuracy loss. Furthermore, efficient hardware architectures for accelerating the compressed LSTM are proposed, which support the inference of multi-layer and multiple time steps. The computation process is judiciously reorganized and the memory access pattern is well optimized, which alleviate the limited memory bandwidth bottleneck and enable higher throughput. Moreover, the parallel processing strategy is carefully designed to make full use of the sparsity introduced by pruning and clipped gating with high hardware utilization efficiency. Implemented on Intel Arria10 S<inline-formula> <tex-math notation="LaTeX">$\times $ </tex-math></inline-formula>660 FPGA running at 200MHz, the proposed design is able to achieve 1.4–2.2 <inline-formula> <tex-math notation="LaTeX">$\times $ </tex-math></inline-formula> energy efficiency and requires significantly less hardware resources compared with the state-of-the-art LSTM implementations.
面向硬件的长短期内存压缩以实现高效推理
DOI: 10.1109/lsp.2018.2834872
发表时间: 2018-05
影响因子: 3.9
作者:
Zhisheng Wang;Jun Lin;Zongfeng Wang
通讯作者: Zongfeng Wang
DOI: --
发表时间: 1993-02
期刊: --
影响因子: --
作者:
J. Garofolo;L. Lamel;W. Fisher;Jonathan G. Fiscus;D. S. Pallett;Nancy L. Dahlgren
通讯作者: J. Garofolo;L. Lamel;W. Fisher;Jonathan G. Fiscus;D. S. Pallett;Nancy L. Dahlgren
DOI: 10.1145/3243176.3243184
发表时间: 2017-11
期刊: Proceedings of the 27th International Conference on Parallel Architectures and Compilation Techniques
影响因子: --
作者:
Franyell Silfa;Gem Dot;J. Arnau;Antonio González
通讯作者: Franyell Silfa;Gem Dot;J. Arnau;Antonio González
DOI: 10.1109/tvlsi.2018.2819190
发表时间: 2018-07
影响因子: 2.8
作者:
Yun Long;Taesik Na;S. Mukhopadhyay
通讯作者: Yun Long;Taesik Na;S. Mukhopadhyay
DOI: --
发表时间: 2016-03
期刊: ArXiv
影响因子: --
作者:
D. Miyashita;Edward H. Lee;B. Murmann
通讯作者: D. Miyashita;Edward H. Lee;B. Murmann