C-LSTM: Enabling Efficient LSTM using Structured Compression Techniques on FPGAs

C-LSTM: Enabling Efficient LSTM using Structured Compression Techniques on FPGAs
复制标题

DOI:
10.1145/3174243.3174253
复制
发表时间:
2018-02
期刊:
Proceedings of the 2018 ACM/SIGDA International Symposium on Field-Programmable Gate Arrays
影响因子:
--
通讯作者:
Shuo Wang;Zhe Li;Caiwen Ding;Bo Yuan;Qinru Qiu;Yanzhi Wang;Yun Liang
Shuo Wang;Zhe Li;Caiwen Ding;Bo Yuan;Qinru Qiu;Yanzhi Wang;Yun Liang
中科院分区:
其他
文献类型:
--
作者:
Shuo Wang;Zhe Li;Caiwen Ding;Bo Yuan;Qinru Qiu;Yanzhi Wang;Yun Liang

文献摘要

被引文献

相似文献

最近,通过增加长短期记忆(LSTM)网络的模型大小,声学识别系统的准确性得到了显着提高。不幸的是,由于片上资源有限,LSTM模型的规模不断增加导致FPGA设计效率低下。先前的工作提出使用基于修剪的压缩技术来减小模型大小,从而加速FPGA上的推理。然而,修剪技术的随机性将模型的密集矩阵转换为高度非结构化的稀疏矩阵,这导致不平衡的计算和不规则的存储器访问,从而损害整体性能和能量效率。相比之下,我们建议使用结构化压缩技术,不仅可以减少LSTM模型的大小,还可以消除计算和内存访问的不规则性。该方法采用块循环矩阵代替稀疏矩阵来压缩权矩阵,并将存储需求从$\mathcalO(k^2)$降低到$\mathcalO(k)$。快速傅立叶变换算法用于进一步加速推理,将计算复杂度从$\mathcalO(k^2)$降低到$\mathcalO(k\textlog k)$。数据路径和激活函数被量化为16位,以提高资源利用率。更重要的是,我们提出了一个名为C-LSTM的综合框架,可以在FPGA上自动优化和实现各种LSTM变体。根据实验结果,在相同的实验设置下,与最先进的LSTM实现相比,C-LSTM在性能和能效方面分别获得了18.8倍和33.5倍的增益,并且精度下降非常小。
Recently, significant accuracy improvement has been achieved for acoustic recognition systems by increasing the model size of Long Short-Term Memory (LSTM) networks. Unfortunately, the ever-increasing size of LSTM model leads to inefficient designs on FPGAs due to the limited on-chip resources. The previous work proposes to use a pruning based compression technique to reduce the model size and thus speedups the inference on FPGAs. However, the random nature of the pruning technique transforms the dense matrices of the model to highly unstructured sparse ones, which leads to unbalanced computation and irregular memory accesses and thus hurts the overall performance and energy efficiency. In contrast, we propose to use a structured compression technique which could not only reduce the LSTM model size but also eliminate the irregularities of computation and memory accesses. This approach employs block-circulant instead of sparse matrices to compress weight matrices and reduces the storage requirement from $\mathcalO (k^2)$ to $\mathcalO (k)$. Fast Fourier Transform algorithm is utilized to further accelerate the inference by reducing the computational complexity from $\mathcalO (k^2)$ to $\mathcalO (k\textlog k)$. The datapath and activation functions are quantized as 16-bit to improve the resource utilization. More importantly, we propose a comprehensive framework called C-LSTM to automatically optimize and implement a wide range of LSTM variants on FPGAs. According to the experimental results, C-LSTM achieves up to 18.8X and 33.5X gains for performance and energy efficiency compared with the state-of-the-art LSTM implementation under the same experimental setup, and the accuracy degradation is very small.