ScalaBFS: A Scalable BFS Accelerator on FPGA-HBM Platform

ScalaBFS: A Scalable BFS Accelerator on FPGA-HBM Platform
复制标题

ScalaBFS:FPGA-HBM 平台上的可扩展 BFS 加速器

DOI:
--
复制
发表时间:
2021
期刊:
Symposium on Field Programmable Gate Arrays
影响因子:
--
通讯作者:
Hai Jin
Hai Jin
中科院分区:
--
文献类型:
--
作者:
Chenhao Liu;Zhiyuan Shao;Kexin Li;Minkang Wu;Jiajie Chen;Ruoshi Li;Xiaofei Liao;Hai Jin

文献摘要

被引文献

相似文献

高带宽内存(HBM)通过将多个存储通道公开到处理单元,从而提供大量的聚合存储器带宽。为了实现高性能,在配置HBM(即FPGA-HBM平台)的FPGA之上构建的加速器需要根据可用的内存频道扩展其性能。在本文中,我们提出了一个用于BFS的加速器(广度优先搜索),称为ScalAbfs,该加速器将内存从处理中访问,以使用可用的HBM内存频道来扩展其性能。此外,通过使用多个处理元素配置每个HBM内存通道,ScalAbf可以充分利用HBM的内存带宽。我们在Xilinx Alveo U280数据中心加速器卡(实际硬件)上实现了ScalABF的原型系统,并在现实世界和无合成规模的图中进行BFS。实验结果表明,根据U280的HBM2子系统的可用内存伪通道(PC),ScalABF几乎可以线性地缩放其性能。通过在U280上完全使用32个PC和建筑物64个处理元件(PES),ScalaBFS可实现高达19.7 GTEPS的性能(GIGA每秒遍历边缘)。当在稀疏现实世界图中进行BFS时,ScalAbfs将等效的GTEP达到了在最先进的NVIDIA V100 GPU上运行的Gun​​rock,该GPU具有64-PC HBM2(内存带宽是U280的两倍)。
High Bandwidth Memory (HBM) provides massive aggregated memory bandwidth by exposing multiple memory channels to the processing units. To achieve high performance, an accelerator built on top of an FPGA configured with HBM (i.e., FPGA-HBM platform) needs to scale its performance according to the available memory channels. In this paper, we propose an accelerator for BFS (Breadth-First Search), named as ScalaBFS, which decouples memory accessing from processing to scale its performance with available HBM memory channels. Moreover, by configuring each HBM memory channel with multiple processing elements, ScalaBFS sufficiently exploits the memory bandwidth of HBM. We implement the prototype system of ScalaBFS and conduct BFS in both real-world and synthetic scale-free graphs on Xilinx Alveo U280 Data Center Accelerator card (real hardware). The experimental results show that ScalaBFS scales its performance almost linearly according to the available memory pseudo channels (PCs) from the HBM2 subsystem of U280. By fully using the 32 PCs and building 64 processing elements (PEs) on U280, ScalaBFS achieves a performance up to 19.7 GTEPS (Giga Traversed Edges Per Second). When conducting BFS in sparse real-world graphs, ScalaBFS achieves equivalent GTEPS to Gunrock running on the state-of-art Nvidia V100 GPU that features 64-PC HBM2 (twice memory bandwidth than U280).