POLAR: A Pipelined/Overlapped FPGA-Based LSTM Accelerator

POLAR: A Pipelined/Overlapped FPGA-Based LSTM Accelerator
复制标题

POLAR:基于 FPGA 的流水线/重叠 LSTM 加速器

DOI:
10.1109/tvlsi.2019.2947639
复制
发表时间:
2020
影响因子:
2.8
通讯作者:
M. Pedram
M. Pedram
中科院分区:
工程技术2区
文献类型:
--
作者:
Erfan Bank;Seyed Abolfazl Ghasemzadeh;M. Kamal;A. Afzali;M. Pedram

文献摘要

被引文献

相似文献

在这篇简报中,提出了一种基于低资源利用率现场可编程门阵列(FPGA)的长短期记忆(LSTM)网络架构,用于加速推理阶段。该架构具有低功耗和高速的特点,通过重叠的操作和流水线的数据路径的时间来实现。此外,该架构需要可忽略的内部存储器大小来存储中间数据,从而导致低资源利用率和简单的路由,这提供了较低的互连延迟(较高的操作频率)。设计者可以通过调整并行化的量来在寄存器传输级(RTL)设计处容易地调整所提出的架构的资源利用率(以及延迟)。这使得将架构映射到不同类型的FPGA的过程变得简单,并受到定义的约束。通过在不同类型的FPGA上实现LSTM网络来评估所提出的架构的有效性。与最近的工作相比,所提出的架构提供了高达约<inline-formula><tex-math notation="LaTeX">$1.6\times $</tex-math></inline-formula>,<inline-formula><tex-math notation="LaTeX">$43.6\times $</tex-math></inline-formula>,<inline-formula><tex-math notation="LaTeX">$21.9\times $</tex-math></inline-formula>,和<inline-formula><tex-math notation="LaTeX">$114.5\times $的</tex-math></inline-formula>频率,功率效率,GOP/s,和GOP/s/W,分别提高。最后,我们提出的架构运行在17.64 GOP/s,这是<inline-formula><tex-math notation="LaTeX">2.31\times $</tex-math></inline-formula>比以前报道的最好的结果快。
In this brief, a low resource utilization field-programmable gate array (FPGA)-based long short-term memory (LSTM) network architecture for accelerating the inference phase is presented. The architecture has low-power and high-speed features that are achieved through overlapping the timing of the operations and pipelining the datapath. Moreover, this architecture requires negligible internal memory size for storing the intermediate data leading to low resource utilization and simple routing, which provides lower interconnect delay (higher operating frequency). A designer may adjust the resource utilization (as well as the latency) of the proposed architecture readily at the register-transfer level (RTL) design by adjusting the amount of parallelization. This makes the process of mapping the architecture onto different types of FPGAs, subject to defined constraints, a simple one. The efficacy of the proposed architecture is assessed by implementing an LSTM network on different types of FPGAs. Compared with the recent works, the proposed architecture provides up to about <inline-formula> <tex-math notation="LaTeX">$1.6\times $ </tex-math></inline-formula>, <inline-formula> <tex-math notation="LaTeX">$43.6\times $ </tex-math></inline-formula>, <inline-formula> <tex-math notation="LaTeX">$21.9\times $ </tex-math></inline-formula>, and <inline-formula> <tex-math notation="LaTeX">$114.5\times $ </tex-math></inline-formula> improvements in frequency, power efficiency, GOP/s, and GOP/s/W, respectively. Finally, our proposed architecture operates at 17.64 GOP/s, which is <inline-formula> <tex-math notation="LaTeX">$2.31\times $ </tex-math></inline-formula> faster than the best previously reported results.