MariusGNN: Resource-Efficient Out-of-Core Training of Graph Neural Networks

MariusGNN: Resource-Efficient Out-of-Core Training of Graph Neural Networks
复制标题

DOI:
10.1145/3552326.3567501
复制
发表时间:
2022-02
期刊:
Proceedings of the Eighteenth European Conference on Computer Systems
影响因子:
--
通讯作者:
R. Waleffe;J. Mohoney;Theodoros Rekatsinas;S. Venkataraman
R. Waleffe;J. Mohoney;Theodoros Rekatsinas;S. Venkataraman
中科院分区:
其他
文献类型:
--
作者:
R. Waleffe;J. Mohoney;Theodoros Rekatsinas;S. Venkataraman

文献摘要

相似文献

我们研究了大规模图的图神经网络(GNNs)的训练。我们重新审视了对十亿级图使用分布式训练的前提,并表明对于适合主内存或单机SSD的图,使用单个GPU的核外流水线训练可以优于最先进的(SoTA)多GPU解决方案。我们介绍MariusGNN,第一个利用整个存储层次结构(包括磁盘)进行GNN训练的系统。MariusGNN引入了一系列数据组织和算法贡献,1)最大限度地减少训练所需的端到端时间,2)确保通过基于磁盘的训练学习的模型表现出与在内存中完全训练的模型相似的准确性。我们评估了MariusGNN与SoTA系统在学习GNN模型方面的对比,发现MariusGNN中的单GPU训练比这些系统中的多GPU训练快8倍,从而实现了数量级的货币成本降低。MariusGNN在www.marius-project.org上开源。
We study training of Graph Neural Networks (GNNs) for large-scale graphs. We revisit the premise of using distributed training for billion-scale graphs and show that for graphs that fit in main memory or the SSD of a single machine, out-of-core pipelined training with a single GPU can outperform state-of-the-art (SoTA) multi-GPU solutions. We introduce MariusGNN, the first system that utilizes the entire storage hierarchy---including disk---for GNN training. MariusGNN introduces a series of data organization and algorithmic contributions that 1) minimize the end-to-end time required for training and 2) ensure that models learned with disk-based training exhibit accuracy similar to those fully trained in memory. We evaluate MariusGNN against SoTA systems for learning GNN models and find that single-GPU training in MariusGNN achieves the same level of accuracy up to 8× faster than multi-GPU training in these systems, thus, introducing an order of magnitude monetary cost reduction. MariusGNN is open-sourced at www.marius-project.org.