Dynasparse: Accelerating GNN Inference through Dynamic Sparsity Exploitation

Dynasparse: Accelerating GNN Inference through Dynamic Sparsity Exploitation
复制标题

DOI:
10.1109/ipdps54959.2023.00032
复制
发表时间:
2023-03
期刊:
2023 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Bingyi Zhang;V. Prasanna
Bingyi Zhang;V. Prasanna
中科院分区:
其他
文献类型:
--
作者:
Bingyi Zhang;V. Prasanna

文献摘要

相似文献

图形神经网络(GNN)推理用于GNN推理中的许多现实应用程序。增加了GNN的数据稀疏性的模型压缩。我们提出了Dynasparse,这是FPGA上的全面硬件软件代码,以通过动态稀疏性加速GNN推断为此,我们将GNN计算内核从基本计算原始图中解脱出来,并探索硬件 - 软件代码,如下所示:1)硬件设计:我们在FPGA上提出了一个新颖的统一加速器设计,以有效地执行各种计算定制的软处理器与加速器紧密结合以执行运行时系统。进行即时数据格式转换,以准备各种计算原语的输入数据;在最先进的FPGA平台Xilinx Alveo U250上实现Dynasparse,并使用广泛使用的GNN型号(GCN,GraphSage,Gin和Sgc)评估设计GNN模型和各种输入图,提出的加速器和动态核对映射与在最新的GNN加速器中进行的静态映射策略相比,平均将推理潜伏期减少了3.73倍与最先进的FPGA实现相比,Dynasparse的状态 - ART CPU(GPU)实现最多可达到56.9×(2.37×)的速度。在加速器执行延迟中达到2.7×加速。
Graph Neural Network (GNN) inference is used in many real-world applications. Data sparsity in GNN inference, including sparsity in the input graph and the GNN model, offer opportunities to further speed up inference. Also, many pruning techniques have been proposed for model compression that increase the data sparsity of GNNs.We propose Dynasparse, a comprehensive hardware-software codesign on FPGA to accelerate GNN inference through dynamic sparsity exploitation. For this, we decouple the GNN computation kernels from the basic computation primitives, and explore hardware-software codesign as follows: 1) Hardware design: We propose a novel unified accelerator design on FPGA to efficiently execute various computation primitives. We develop a customized soft processor that is tightly coupled with the accelerator to execute a runtime system. Moreover, we develop efficient hardware mechanisms to profile the data sparsity and perform on-the-fly data format transformation to prepare the input data for various computation primitives; 2) Software design: We develop a runtime system that works synergistically with the accelerator to perform dynamic kernel-to-primitive mapping based on data sparsity. We implement Dynasparse on a state-of-the-art FPGA platform, Xilinx Alveo U250, and evaluate the design using widely used GNN models (GCN, GraphSAGE, GIN and SGC). For the above GNN models and various input graphs, the proposed accelerator and dynamic kernel-to-primitive mapping reduces the inference latency by 3.73× on the average compared with the static mapping strategies employed in the state-of-the-art GNN accelerators. Compared with state-of-the-art CPU (GPU) implementations, Dynasparse achieves up to 56.9× (2.37×) speedup in end-to-end latency. Compared with state-of-the-art FPGA implementations, Dynasparse achieves 2.7× speedup in accelerator execution latency.