MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU Platforms

MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU Platforms
复制标题

DOI:
--
复制
发表时间:
2022-09
期刊:
--
影响因子:
--
通讯作者:
Yuke Wang;Boyuan Feng;Zheng Wang;Tong Geng;K. Barker;Ang Li;Yufei Ding
Yuke Wang;Boyuan Feng;Zheng Wang;Tong Geng;K. Barker;Ang Li;Yufei Ding
中科院分区:
其他
文献类型:
--
作者:
Yuke Wang;Boyuan Feng;Zheng Wang;Tong Geng;K. Barker;Ang Li;Yufei Ding

文献摘要

相似文献

图神经网络(GNN)输入图的大小不断增加,凸显了使用多GPU平台的需求。然而,现有的多GPU GNN系统基于扩展密集DNN的传统实践来单独优化计算和通信。对于不规则稀疏和细粒度的GNN工作负载,这样的解决方案错过了联合调度/优化计算和通信操作以实现高性能交付的机会。为此,我们提出了MGG,这是一种新的系统设计,可以在多GPU平台上加速全图GNN。MGG的核心是其新颖的动态软件流水线,以促进GPU内核内的细粒度计算通信重叠。具体而言,MGG引入了GNN定制的流水线构造和GPU感知的流水线映射,以促进工作负载平衡和操作重叠。MGG还结合了智能运行时设计与分析建模和优化算法,以动态提高执行性能。广泛的评估表明,MGG在各种设置中的性能优于最先进的全图GNN系统:平均分别比DGL,MGG-UVM和ROC快4.41倍,4.81倍和10.83倍。
The increasing size of input graphs for graph neural networks (GNNs) highlights the demand for using multi-GPU platforms. However, existing multi-GPU GNN systems optimize the computation and communication individually based on the conventional practice of scaling dense DNNs. For irregularly sparse and fine-grained GNN workloads, such solutions miss the opportunity to jointly schedule/optimize the computation and communication operations for high-performance delivery. To this end, we propose MGG, a novel system design to accelerate full-graph GNNs on multi-GPU platforms. The core of MGG is its novel dynamic software pipeline to facilitate fine-grained computation-communication overlapping within a GPU kernel. Specifically, MGG introduces GNN-tailored pipeline construction and GPU-aware pipeline mapping to facilitate workload balancing and operation overlapping. MGG also incorporates an intelligent runtime design with analytical modeling and optimization heuristics to dynamically improve the execution performance. Extensive evaluation reveals that MGG outperforms state-of-the-art full-graph GNN systems across various settings: on average 4.41X, 4.81X, and 10.83X faster than DGL, MGG-UVM, and ROC, respectively.