GraphLily: Accelerating Graph Linear Algebra on HBM-Equipped FPGAs

GraphLily: Accelerating Graph Linear Algebra on HBM-Equipped FPGAs
复制标题

DOI:
10.1109/iccad51958.2021.9643582
复制
发表时间:
2021-11
期刊:
2021 IEEE/ACM International Conference On Computer Aided Design (ICCAD)
影响因子:
--
通讯作者:
Yuwei Hu;Yixiao Du;Ecenur Ustun;Zhiru Zhang
Yuwei Hu;Yixiao Du;Ecenur Ustun;Zhiru Zhang
中科院分区:
其他
文献类型:
--
作者:
Yuwei Hu;Yixiao Du;Ecenur Ustun;Zhiru Zhang

文献摘要

相似文献

图形处理通常是由于较低的计算与内存访问率和不规则数据访问模式而受到的内存绑定。新兴的高带宽内存(HBM)通过提供可以同时服务内存请求的多个通道来提供出色的带宽,从而带来了显着提高图形处理性能的潜力。本文提出了图形线性代数覆盖图,以加速配备HBM的FPGA上的图形处理。 GraphLily通过采用Graphblas编程接口来支持一组丰富的图形算法,该界面将图形算法作为稀疏线性代数操作制定。 GraphLily为Graphblas中的两个广泛使用的核提供了有效的,内存优化的加速器,即稀疏 - 矩阵密集矢量乘法(SPMV)和稀疏的Matrix Sparse-Vector-vector乘数(SPMSPV)。 SPMV加速器使用针对HBM量身定制的稀疏矩阵存储格式,该格式可以启用流媒体,矢量化访问到每个通道,并同时访问多个通道。此外,SPMV加速器通过引入可扩展的片上缓冲区设计来利用密集矢量访问的数据。 SPMSPV加速器补充了SPMV加速器,以处理输入向量具有较高稀疏性的情况。 Graphlilly进一步构建了中间件以提供运行时支持。借助此中间件,我们可以将现有的Graphblas程序移到FPGA中,并对旨在CPU/GPU执行的原始代码进行一些修改。评估表明,与CPU和GPU上的最新图形处理框架相比,Graphlily的达到高达2.5 x和1.1 x的吞吐量,同时将能耗降低了8.1 x和2.4 x;与fPGA上的先前单用途的图形加速器相比,Graphlily实现了1.2 x -1.9 x较高的吞吐量。
Graph processing is typically memory bound due to low compute to memory access ratio and irregular data access pattern. The emerging high-bandwidth memory (HBM) delivers exceptional bandwidth by providing multiple channels that can service memory requests concurrently, thus bringing the potential to significantly boost the performance of graph processing. This paper proposes GraphLily, a graph linear algebra overlay, to accelerate graph processing on HBM-equipped FPGAs. GraphLily supports a rich set of graph algorithms by adopting the GraphBLAS programming interface, which formulates graph algorithms as sparse linear algebra operations. GraphLily provides efficient, memory-optimized accelerators for the two widely-used kernels in GraphBLAS, namely, sparse-matrix dense-vector multiplication (SpMV) and sparse-matrix sparse-vector multiplication (SpMSpV). The SpMV accelerator uses a sparse matrix storage format tailored to HBM that enables streaming, vectorized accesses to each channel and concurrent accesses to multiple channels. Besides, the SpMV accelerator exploits data reuse in accesses of the dense vector by introducing a scalable on-chip buffer design. The SpMSpV accelerator complements the SpMV accelerator to handle cases where the input vector has a high sparsity. GraphLily further builds a middleware to provide runtime support. With this middleware, we can port existing GraphBLAS programs to FPGAs with slight modifications to the original code intended for CPU/GPU execution. Evaluation shows that compared with state-of-the-art graph processing frameworks on CPUs and GPUs, GraphLily achieves up to 2.5 x and 1.1 x higher throughput, while reducing the energy consumption by 8.1 x and 2.4 x; compared with prior single-purpose graph accelerators on FPGAs, GraphLily achieves 1.2 x -1.9 x higher throughput.