PipeGCN: Efficient Full-Graph Training of Graph Convolutional Networks with Pipelined Feature Communication

PipeGCN: Efficient Full-Graph Training of Graph Convolutional Networks with Pipelined Feature Communication
复制标题

DOI:
10.48550/arxiv.2203.10428
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Cheng Wan;Youjie Li;Cameron R. Wolfe;Anastasios Kyrillidis;Namjae Kim;Yingyan Lin
Cheng Wan;Youjie Li;Cameron R. Wolfe;Anastasios Kyrillidis;Namjae Kim;Yingyan Lin
中科院分区:
其他
文献类型:
--
作者:
Cheng Wan;Youjie Li;Cameron R. Wolfe;Anastasios Kyrillidis;Namjae Kim;Yingyan Lin

文献摘要

相似文献

图卷积网络(GCNs)是学习图结构数据的最先进方法,训练大规模GCN需要在多个加速器上进行分布式训练,以便每个加速器能够容纳一个划分后的子图。然而,分布式GCN训练在每次训练迭代期间,每个GCN层的分区之间传递节点特征和特征梯度会产生高昂的开销,这限制了可达到的训练效率和模型可扩展性。为此,我们提出了PipeGCN,这是一种简单而有效的方案,它通过将分区间通信与分区内计算进行流水线操作来隐藏通信开销。对于高效的GCN训练进行流水线操作并非易事,因为传递的节点特征/梯度会变得过时,从而可能损害收敛性,抵消流水线的优势。值得注意的是,对于同时存在过时特征和过时特征梯度的GCN训练的收敛速度,人们知之甚少。这项工作不仅提供了理论收敛分析,还发现PipeGCN的收敛速度接近没有任何过时性的普通分布式GCN训练的收敛速度。此外,我们开发了一种平滑方法来进一步提高PipeGCN的收敛性。大量实验表明,PipeGCN可以大幅提高训练吞吐量(1.7倍~28.5倍),同时达到与其普通对应方法以及现有全图训练方法相同的精度。代码可在https://github.com/RICE - EIC/PipeGCN获取。
Graph Convolutional Networks (GCNs) is the state-of-the-art method for learning graph-structured data, and training large-scale GCNs requires distributed training across multiple accelerators such that each accelerator is able to hold a partitioned subgraph. However, distributed GCN training incurs prohibitive overhead of communicating node features and feature gradients among partitions for every GCN layer during each training iteration, limiting the achievable training efficiency and model scalability. To this end, we propose PipeGCN, a simple yet effective scheme that hides the communication overhead by pipelining inter-partition communication with intra-partition computation. It is non-trivial to pipeline for efficient GCN training, as communicated node features/gradients will become stale and thus can harm the convergence, negating the pipeline benefit. Notably, little is known regarding the convergence rate of GCN training with both stale features and stale feature gradients. This work not only provides a theoretical convergence analysis but also finds the convergence rate of PipeGCN to be close to that of the vanilla distributed GCN training without any staleness. Furthermore, we develop a smoothing method to further improve PipeGCN's convergence. Extensive experiments show that PipeGCN can largely boost the training throughput (1.7x~28.5x) while achieving the same accuracy as its vanilla counterpart and existing full-graph training methods. The code is available at https://github.com/RICE-EIC/PipeGCN.