POLAR: A Pipelined/Overlapped FPGA-Based LSTM Accelerator
POLAR: A Pipelined/Overlapped FPGA-Based LSTM Accelerator
复制标题
POLAR:基于 FPGA 的流水线/重叠 LSTM 加速器
DOI:
10.1109/tvlsi.2019.2947639
复制
发表时间:
2020
影响因子:
2.8
通讯作者:
M. Pedram
中科院分区:
文献类型:
--
作者:
Erfan Bank;Seyed Abolfazl Ghasemzadeh;M. Kamal;A. Afzali;M. Pedram
In this brief, a low resource utilization field-programmable gate array (FPGA)-based long short-term memory (LSTM) network architecture for accelerating the inference phase is presented. The architecture has low-power and high-speed features that are achieved through overlapping the timing of the operations and pipelining the datapath. Moreover, this architecture requires negligible internal memory size for storing the intermediate data leading to low resource utilization and simple routing, which provides lower interconnect delay (higher operating frequency). A designer may adjust the resource utilization (as well as the latency) of the proposed architecture readily at the register-transfer level (RTL) design by adjusting the amount of parallelization. This makes the process of mapping the architecture onto different types of FPGAs, subject to defined constraints, a simple one. The efficacy of the proposed architecture is assessed by implementing an LSTM network on different types of FPGAs. Compared with the recent works, the proposed architecture provides up to about <inline-formula> <tex-math notation="LaTeX">$1.6\times $ </tex-math></inline-formula>, <inline-formula> <tex-math notation="LaTeX">$43.6\times $ </tex-math></inline-formula>, <inline-formula> <tex-math notation="LaTeX">$21.9\times $ </tex-math></inline-formula>, and <inline-formula> <tex-math notation="LaTeX">$114.5\times $ </tex-math></inline-formula> improvements in frequency, power efficiency, GOP/s, and GOP/s/W, respectively. Finally, our proposed architecture operates at 17.64 GOP/s, which is <inline-formula> <tex-math notation="LaTeX">$2.31\times $ </tex-math></inline-formula> faster than the best previously reported results.