BNS-GCN: Efficient Full-Graph Training of Graph Convolutional Networks with Partition-Parallelism and Random Boundary Node Sampling

BNS-GCN: Efficient Full-Graph Training of Graph Convolutional Networks with Partition-Parallelism and Random Boundary Node Sampling
复制标题

DOI:
10.48550/arxiv.2203.10983
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Cheng Wan;Youjie Li;Ang Li;Namjae Kim;Yingyan Lin
Cheng Wan;Youjie Li;Ang Li;Namjae Kim;Yingyan Lin
中科院分区:
其他
文献类型:
--
作者:
Cheng Wan;Youjie Li;Ang Li;Namjae Kim;Yingyan Lin

文献摘要

相似文献

图卷积网络(GCN)已成为基于图的学​​习任务的最先进方法。然而,大规模训练 GCN 仍然具有挑战性,阻碍了对更复杂的 GCN 架构的探索及其在现实世界大图中的应用。虽然考虑图分区和分布式训练来应对这一挑战可能是很自然的,但由于现有设计的限制,这个方向在之前的工作中只触及了表面。在这项工作中,我们首先分析了为什么分布式 GCN 训练无效,并确定其根本原因是每个分区子图的边界节点数量过多,这很容易导致 GCN 训练的内存和通信成本爆炸。此外,我们提出了一种简单而有效的方法,称为 BNS-GCN,它采用随机边界节点采样来实现高效且可扩展的分布式 GCN 训练。实验和消融研究一致验证了 BNS-GCN 的有效性,例如,将吞吐量提高高达 16.2 倍,将内存使用量降低高达 58%,同时保持全图精度。此外,理论和实证分析都表明,BNS-GCN 比现有的基于采样的方法具有更好的收敛性。我们相信,我们的 BNS-GCN 为大规模 GCN 训练开辟了新的范式。代码可在 https://github.com/RICE-EIC/BNS-GCN 获取。
Graph Convolutional Networks (GCNs) have emerged as the state-of-the-art method for graph-based learning tasks. However, training GCNs at scale is still challenging, hindering both the exploration of more sophisticated GCN architectures and their applications to real-world large graphs. While it might be natural to consider graph partition and distributed training for tackling this challenge, this direction has only been slightly scratched the surface in the previous works due to the limitations of existing designs. In this work, we first analyze why distributed GCN training is ineffective and identify the underlying cause to be the excessive number of boundary nodes of each partitioned subgraph, which easily explodes the memory and communication costs for GCN training. Furthermore, we propose a simple yet effective method dubbed BNS-GCN that adopts random Boundary-Node-Sampling to enable efficient and scalable distributed GCN training. Experiments and ablation studies consistently validate the effectiveness of BNS-GCN, e.g., boosting the throughput by up to 16.2x and reducing the memory usage by up to 58%, while maintaining a full-graph accuracy. Furthermore, both theoretical and empirical analysis show that BNS-GCN enjoys a better convergence than existing sampling-based methods. We believe that our BNS-GCN has opened up a new paradigm for enabling GCN training at scale. The code is available at https://github.com/RICE-EIC/BNS-GCN.