Compute-Capable Block RAMs for Efficient Deep Learning Acceleration on FPGAs

Compute-Capable Block RAMs for Efficient Deep Learning Acceleration on FPGAs
复制标题

DOI:
10.1109/fccm51124.2021.00018
复制
发表时间:
2021-05
期刊:
2021 IEEE 29th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)
影响因子:
--
通讯作者:
Xiaowei Wang;Vidushi Goyal;Jiecao Yu;V. Bertacco;Andrew Boutros;Eriko Nurvitadhi;C. Augustine;R. Iyer;R. Das
Xiaowei Wang;Vidushi Goyal;Jiecao Yu;V. Bertacco;Andrew Boutros;Eriko Nurvitadhi;C. Augustine;R. Iyer;R. Das
中科院分区:
其他
文献类型:
--
作者:
Xiaowei Wang;Vidushi Goyal;Jiecao Yu;V. Bertacco;Andrew Boutros;Eriko Nurvitadhi;C. Augustine;R. Iyer;R. Das

文献摘要

被引文献

相似文献

FPGA片上记忆的密度一直在不断增加,而现代FPGA在其可重新配置的结构上分布着数千个块RAM(BRAM)。 . In this work, we propose enhancing the ubiquitous FPGA BRAMs with in-memory compute-capabilities. As a result, BRAMs can act as normal storage units or their bitlines can be re-purposed as SIMD lanes executing bit-serial arithmetic operations. Our拟议的建筑变化导致1.6×和2.3×增加峰值多重蓄能的大型Stratix 10 FPGA的吞吐量,FPGA模具尺寸的最低成本仅增加1.8%,而Bram的界面与可编程路由没有变化然后,我们提出RIMA,RIMA可重新配置的内存加速器体系结构(DL)推理。 ART Brainwave DL软处理器用于8位整数和块浮点精度,此外,在Stratix 10 FPGA上实施的RIMA通过具有计算能力的BRAM可以达到相同的数量级, GPU一代。
The density of FPGA on-chip memory has been continuously increasing with modern FPGAs having thousands of block RAMs (BRAMs) distributed across their reconfigurable fabric. These distributed BRAMs can provide a tremendous amount of on-chip bandwidth for efficient acceleration of data-intensive applications. In this work, we propose enhancing the ubiquitous FPGA BRAMs with in-memory compute-capabilities. As a result, BRAMs can act as normal storage units or their bitlines can be re-purposed as SIMD lanes executing bit-serial arithmetic operations. Our proposed architectural change results in 1.6× and 2.3× increase in the peak multiply-accumulate throughput of a large Stratix 10 FPGA, at a minimal cost of only 1.8% increase in the FPGA die size and no change to the BRAM’s interface to the programmable routing. Then, we present RIMA, a reconfigurable in-memory accelerator architecture for deep learning (DL) inference. RIMA exploits the proposed compute-capable BRAMs and the FPGA’s reconfigurability to achieve 1.25× and 3× higher performance compared to the state-of-the-art Brainwave DL soft processor for 8-bit integer and block floating-point precisions, respectively. In addition, RIMA implemented on a Stratix 10 FPGA enhanced with compute-capable BRAMs can achieve an order of magnitude higher performance compared to a same-generation GPU.