Energy-Efficient LSTM Inference Accelerator for Real-Time Causal Prediction

Energy-Efficient LSTM Inference Accelerator for Real-Time Causal Prediction
复制标题

用于实时因果预测的节能 LSTM 推理加速器

DOI:
10.1145/3495006
复制
发表时间:
2022
影响因子:
1.4
通讯作者:
Cong, Jason
Cong, Jason
中科院分区:
计算机科学4区
文献类型:
--
作者:
Chen, Zhe;Blair, Hugh T.;Cong, Jason

文献摘要

相似文献

不断增长的边缘应用程序通常需要较短的处理延迟和较高的能源效率,以满足严格的时序和功率预算。在这项工作中,我们提出紧凑的长短期记忆(LSTM)模型可以近似传统的非因果算法,减少延迟并提高实时因果预测的效率,特别是对于闭环反馈应用中的神经信号处理。我们利用细粒度并行性以及流水线前馈和循环更新来设计 LSTM 推理加速器。我们还提出了一种位稀疏量化方法,该方法可以通过用移位运算符代替乘法器来减少电路面积和功耗。我们探索了修剪和量化方法的不同组合,以便对从脑电图 (EEG) 和钙图像处理应用程序收集的数据集进行节能的 LSTM 推理。评估结果表明,我们提出的 LSTM 推理加速器可以实现 1.19 GOPS/mW 的能效。与 16 位基线实现相比,具有 2 sbit/16 位稀疏量化和 60% 稀疏度的 LSTM 加速器可分别减少 54.1% 和 56.3% 的电路面积和功耗。
Ever-growing edge applications often require short processing latency and high energy efficiency to meet strict timing and power budget. In this work, we propose that the compact long short-term memory (LSTM) model can approximate conventionalacausalalgorithms with reduced latency and improved efficiency for real-time causal prediction, especially for the neural signal processing in closed-loop feedback applications. We design an LSTM inference accelerator by taking advantage of the fine-grained parallelism and pipelined feedforward and recurrent updates. We also propose a bit-sparse quantization method that can reduce the circuit area and power consumption by replacing the multipliers with the bit-shift operators. We explore different combinations of pruning and quantization methods for energy-efficient LSTM inference on datasets collected from the electroencephalogram (EEG) and calcium image processing applications. Evaluation results show that our proposed LSTM inference accelerator can achieve 1.19 GOPS/mW energy efficiency. The LSTM accelerator with 2-sbit/16-bit sparse quantization and 60% sparsity can reduce the circuit area and power consumption by 54.1% and 56.3%, respectively, compared with a 16-bit baseline implementation.