BoostGCN: A Framework for Optimizing GCN Inference on FPGA

BoostGCN: A Framework for Optimizing GCN Inference on FPGA
复制标题

DOI:
10.1109/fccm51124.2021.00012
复制
发表时间:
2021-02
期刊:
2021 IEEE 29th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM)
影响因子:
--
通讯作者:
Bingyi Zhang;R. Kannan;V. Prasanna
Bingyi Zhang;R. Kannan;V. Prasanna
中科院分区:
其他
文献类型:
--
作者:
Bingyi Zhang;R. Kannan;V. Prasanna

文献摘要

被引文献

相似文献

图卷积网络(GCNS)给推荐系统、流量预测等大数据应用带来了革命性的变化。然而,由于(1)巨大的外部存储流量和不规则的存储器访问,(2)由于度分布的偏斜导致的负载不平衡,以及(3)由于算法的两个不同的计算阶段导致的级内负载不平衡,加速GCN推理是具有挑战性的。针对上述问题,我们提出了一种基于现场可编程门阵列的GCN推理优化框架BoostGCN。首先,我们开发了一种新的硬件感知的以划分为中心的特征聚集(PCFA)方案,该方案将三维划分与以顶点为中心的计算范式相结合。这提高了片上数据的重用性,并减少了与外部存储器的总数据通信量。其次,我们设计了一种新的硬件体系结构,使两个不同计算阶段的流水线执行成为可能。我们提出了一种低开销的任务调度策略,以减少这两个计算阶段造成的流水线停顿。第三,利用优化的RTL模板,在FPGA上提供了一个完整的GCN加速框架。它可以根据定制的配置生成硬件设计,并适用于各种GCN模型。使用我们的框架,我们在最先进的FPGA平台上为各种GCN模型生成加速器,并使用广泛使用的数据集来评估我们的设计。实验结果表明,与目前最先进的处理器(≈100x)、图形处理器(≈30x)、先前的fpga加速器(3-45x)相比,该框架产生的加速器具有显著的加速比。
Graph convolutional networks (GCNs) have revolutionized many big data applications, such as recommendation systems, traffic prediction, etc. However, accelerating GCN inference is challenging due to (1) massive external memory traffic and irregular memory access, (2) workload imbalance due to skewed degree distribution, and (3) intra-stage load imbalance caused by two heterogeneous computation phases of the algorithm. To address the above challenges, we propose a framework named BoostGCN to optimize GCN inference on FPGA. First, we develop a novel hardware-aware Partition-Centric Feature Aggregation (PCFA) scheme that leverages 3-D partitioning with the vertex-centric computing paradigm. This increases on-chip data reuse and reduces the total data communication volume with external memory. Second, we design a novel hardware architecture to enable pipelined execution of the two heterogeneous computation phases. We develop a low-overhead task scheduling strategy to reduce the pipeline stalls caused by the two computation phases. Third, we provide a complete GCN acceleration framework on FPGA with optimized RTL templates. It can generate hardware designs based on the customized configuration and is adaptable to various GCN models. Using our framework, we generate accelerators for various GCN models on a state-of-the-art FPGA platform and evaluate our designs using widely used datasets. Experimental results show that the accelerators produced by our framework achieve significant speedup compared with state-of-the-art implementations on CPU (≈ 100×), GPU (≈ 30×), prior FPGA accelerator (3-45)×.