CoMeFa: Deploying Compute-in-Memory on FPGAs for Deep Learning Acceleration

CoMeFa: Deploying Compute-in-Memory on FPGAs for Deep Learning Acceleration
复制标题

DOI:
10.1145/3603504
复制
发表时间:
2023-06
影响因子:
2.3
通讯作者:
Aman Arora;Atharva Bhamburkar;Aatman Borda;T. Anand;Rishabh Sehgal;Bagus Hanindhito;P. Gaillardon;J. Kulkarni;L. John
Aman Arora;Atharva Bhamburkar;Aatman Borda;T. Anand;Rishabh Sehgal;Bagus Hanindhito;P. Gaillardon;J. Kulkarni;L. John
中科院分区:
计算机科学3区
文献类型:
--
作者:
Aman Arora;Atharva Bhamburkar;Aatman Borda;T. Anand;Rishabh Sehgal;Bagus Hanindhito;P. Gaillardon;J. Kulkarni;L. John

文献摘要

被引文献

相似文献

块随机访问记忆(BRAMS)是FPGA的储物室,为使用逻辑块和数字信号处理切片实现的计算单元提供了广泛的芯片内存带宽。 FPGAS的块随机访问记忆(RAM) RAMS使用FPGA BRAM的真实双端口性质,并包含多个可配置的单位式处理元素RAMS到FPGAS显着提高了其计算密度,同时还会探索这些RAM的两个体系结构:COMEFA-D (对延迟进行了优化)和comefa-a(对区域进行了优化)特别适用于DL等平行和计算密集型应用程序,但是这些多功能块在信号处理和数据库等。通过增加Intel Arria 10型FPGA(COMEFA-D(COMEFA-A)RAM,以3.8%(1.2%)的面积,并具有算法的改进和有效的映射来自各种应用的微型计算机的2.55×(1.85×),在多个深层神经网络上最多可以替换所有或一些bram的东西。 FPGA中的公羊可以使它们更好地加速DL工作负载。
Block random access memories (BRAMs) are the storage houses of FPGAs, providing extensive on-chip memory bandwidth to the compute units implemented using logic blocks and digital signal processing slices. We propose modifying BRAMs to convert them to CoMeFa (Compute-in-Memory Blocks for FPGAs) random access memories (RAMs). These RAMs provide highly parallel compute-in-memory by combining computation and storage capabilities in one block. CoMeFa RAMs utilize the true dual-port nature of FPGA BRAMs and contain multiple configurable single-bit bit-serial processing elements. CoMeFa RAMs can be used to compute with any precision, which is extremely important for applications like deep learning (DL). Adding CoMeFa RAMs to FPGAs significantly increases their compute density while also reducing data movement. We explore and propose two architectures of these RAMs: CoMeFa-D (optimized for delay) and CoMeFa-A (optimized for area). Compared to existing proposals, CoMeFa RAMs do not require changing the underlying static RAM technology like simultaneously activating multiple wordlines on the same port, and are practical to implement. CoMeFa RAMs are especially suitable for parallel and compute-intensive applications like DL, but these versatile blocks find applications in diverse applications like signal processing and databases, among others. By augmenting an Intel Arria 10–like FPGA with CoMeFa-D (CoMeFa-A) RAMs at the cost of 3.8% (1.2%) area, and with algorithmic improvements and efficient mapping, we observe a geomean speedup of 2.55× (1.85×) across microbenchmarks from various applications and a geomean speedup of up to 2.5× across multiple deep neural networks. Replacing all or some BRAMs with CoMeFa RAMs in FPGAs can make them better accelerators of DL workloads.