Accelerating graph sampling for graph machine learning using GPUs

Accelerating graph sampling for graph machine learning using GPUs
复制标题

DOI:
10.1145/3447786.3456244
复制
发表时间:
2020-09
期刊:
Proceedings of the Sixteenth European Conference on Computer Systems
影响因子:
--
通讯作者:
Abhinav Jangda;Sandeep Polisetty;Arjun Guha;M. Serafini
Abhinav Jangda;Sandeep Polisetty;Arjun Guha;M. Serafini
中科院分区:
其他
文献类型:
--
作者:
Abhinav Jangda;Sandeep Polisetty;Arjun Guha;M. Serafini

文献摘要

被引文献

相似文献

表示学习算法自动学习数据的特征。图数据的几种表示学习算法,如DeepWalk,node 2 vec和Graph-SAGE,对图进行采样以产生适合训练DNN的小批量。然而,采样时间可能是训练时间的很大一部分,并且现有系统不能有效地并行采样。采样是一个“非常并行”的问题,可能会出现GPU加速,但图形的不规则性使得很难有效地使用GPU资源。本文介绍了NextDoor,一个旨在有效地在GPU上执行图形采样的系统。NextDoor采用了一种新的图形采样方法,我们称之为transit-parallelism,它允许负载平衡和边缘缓存。NextDoor为最终用户提供了编写各种图形采样算法的高级抽象。我们实现了几个图形采样应用程序,并表明NextDoor运行它们的数量级比现有的系统更快。
Representation learning algorithms automatically learn the features of data. Several representation learning algorithms for graph data, such as DeepWalk, node2vec, and Graph-SAGE, sample the graph to produce mini-batches that are suitable for training a DNN. However, sampling time can be a significant fraction of training time, and existing systems do not efficiently parallelize sampling. Sampling is an "embarrassingly parallel" problem and may appear to lend itself to GPU acceleration, but the irregularity of graphs makes it hard to use GPU resources effectively. This paper presents NextDoor, a system designed to effectively perform graph sampling on GPUs. NextDoor employs a new approach to graph sampling that we call transit-parallelism, which allows load balancing and caching of edges. NextDoor provides end-users with a high-level abstraction for writing a variety of graph sampling algorithms. We implement several graph sampling applications, and show that NextDoor runs them orders of magnitude faster than existing systems.