E-PUR: an energy-efficient processing unit for recurrent neural networks

E-PUR: an energy-efficient processing unit for recurrent neural networks
复制标题

DOI:
10.1145/3243176.3243184
复制
发表时间:
2017-11
期刊:
Proceedings of the 27th International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
Franyell Silfa;Gem Dot;J. Arnau;Antonio González
Franyell Silfa;Gem Dot;J. Arnau;Antonio González
中科院分区:
其他
文献类型:
--
作者:
Franyell Silfa;Gem Dot;J. Arnau;Antonio González

文献摘要

被引文献

相似文献

递归神经网络(RNNs)是自动语音识别、机器翻译或图像描述等新兴应用的关键技术。长短期记忆(LSTM)网络是最成功的RNN实现,因为它们可以学习长期依赖关系以达到较高的准确性。不幸的是,LSTM网络的循环特性极大地限制了并行性的数量,因此,多核cpu和多核gpu在RNN推理中表现出较差的效率。在本文中,我们提出了E-PUR,一种适合LSTM计算要求的节能处理单元。E-PUR的主要目标是支持用于低功耗移动设备的大型循环神经网络。E-PUR为LSTM网络提供了一种高效的硬件实现,可以灵活地支持各种应用。它的一个主要新颖之处是我们称之为最大权重局部性(MWL)的技术,它改善了获取突触权重的内存访问的时间局部性,在很大程度上减少了内存需求。我们的实验结果表明,E-PUR在不同的LSTM网络上实现了实时性能,同时相对于通用处理器和gpu降低了几个数量级的能耗,并且需要非常小的芯片面积。与现代移动SoC (NVIDIA Tegra X1)相比,E-PUR平均节能88倍。
Recurrent Neural Networks (RNNs) are a key technology for emerging applications such as automatic speech recognition, machine translation or image description. Long Short Term Memory (LSTM) networks are the most successful RNN implementation, as they can learn long term dependencies to achieve high accuracy. Unfortunately, the recurrent nature of LSTM networks significantly constrains the amount of parallelism and, hence, multicore CPUs and many-core GPUs exhibit poor efficiency for RNN inference. In this paper, we present E-PUR, an energy-efficient processing unit tailored to the requirements of LSTM computation. The main goal of E-PUR is to support large recurrent neural networks for low-power mobile devices. E-PUR provides an efficient hardware implementation of LSTM networks that is flexible to support diverse applications. One of its main novelties is a technique that we call Maximizing Weight Locality (MWL), which improves the temporal locality of the memory accesses for fetching the synaptic weights, reducing the memory requirements by a large extent. Our experimental results show that E-PUR achieves real-time performance for different LSTM networks, while reducing energy consumption by orders of magnitude with respect to general-purpose processors and GPUs, and it requires a very small chip area. Compared to a modern mobile SoC, an NVIDIA Tegra X1, E-PUR provides an average energy reduction of 88x.