Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads

Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
John Thorpe;Yifan Qiao;Jon Eyolfson;Shen Teng;Guanzhou Hu;Zhihao Jia;Jinliang Wei;Keval Vora-Keval-Vor
John Thorpe;Yifan Qiao;Jon Eyolfson;Shen Teng;Guanzhou Hu;Zhihao Jia;Jinliang Wei;Keval Vora-Keval-Vor
中科院分区:
其他
文献类型:
--
作者:
John Thorpe;Yifan Qiao;Jon Eyolfson;Shen Teng;Guanzhou Hu;Zhihao Jia;Jinliang Wei;Keval Vora-Keval-Vor

文献摘要

被引文献

相似文献

图神经网络(GNN)实现了对结构化图数据的深度学习。有两个主要的GNN训练障碍:1)它依赖于具有许多gpu的高端服务器,这些服务器的购买和维护成本很高;2)gpu上有限的内存无法扩展到今天的十亿边缘图。本文介绍了Dorylus:一个用于训练GNNs的分布式系统。独特的是,Dorylus可以利用无服务器计算以低成本提高可扩展性。指导我们设计的关键观点是计算分离。计算分离使得构建一个深度的、有界的异步管道成为可能,其中图和张量并行任务可以完全重叠,有效地隐藏了Lambdas带来的网络延迟。在数千个Lambda线程的帮助下,Dorylus将GNN训练扩展到十亿边图。目前,对于大型图形,CPU服务器比GPU服务器提供最好的性价比。仅仅在CPU服务器上使用Lambdas,每美元的性能就比只使用CPU服务器训练高出2.75倍。具体来说,Dorylus在处理大规模稀疏图时比GPU服务器快1.22倍,便宜4.83倍。与现有的基于采样的系统相比,Dorylus的速度快3.8倍,价格便宜10.7倍。
A graph neural network (GNN) enables deep learning on structured graph data. There are two major GNN training obstacles: 1) it relies on high-end servers with many GPUs which are expensive to purchase and maintain, and 2) limited memory on GPUs cannot scale to today's billion-edge graphs. This paper presents Dorylus: a distributed system for training GNNs. Uniquely, Dorylus can take advantage of serverless computing to increase scalability at a low cost. The key insight guiding our design is computation separation. Computation separation makes it possible to construct a deep, bounded-asynchronous pipeline where graph and tensor parallel tasks can fully overlap, effectively hiding the network latency incurred by Lambdas. With the help of thousands of Lambda threads, Dorylus scales GNN training to billion-edge graphs. Currently, for large graphs, CPU servers offer the best performance-per-dollar over GPU servers. Just using Lambdas on top of CPU servers offers up to 2.75x more performance-per-dollar than training only with CPU servers. Concretely, Dorylus is 1.22x faster and 4.83x cheaper than GPU servers for massive sparse graphs. Dorylus is up to 3.8x faster and 10.7x cheaper compared to existing sampling-based systems.