An Edge Attribute-Wise Partitioning and Distributed Processing of R-GCN Using GPUs

An Edge Attribute-Wise Partitioning and Distributed Processing of R-GCN Using GPUs
复制标题

DOI:
10.1007/978-3-030-71593-9_10
复制
发表时间:
2021-02
期刊:
Euro-Par 2020: Parallel Processing Workshops
影响因子:
--
通讯作者:
Tokio Kibata;Mineto Tsukada;Hiroki Matsutani
Tokio Kibata;Mineto Tsukada;Hiroki Matsutani
中科院分区:
其他
文献类型:
--
作者:
Tokio Kibata;Mineto Tsukada;Hiroki Matsutani

文献摘要

相似文献

R-GCN(Relational Graph Convolutional Network)是图神经网络的一种。该模型试图通过考虑图结构数据(如知识库)中的边的方向和类型来预测潜在信息。该模型为每个边属性建立权重矩阵。因此,神经网络的大小随着边缘类型的数量线性增加。虽然GPU可以用于加速R-GCN处理,但权重矩阵的大小可能超过GPU设备内存。为了解决这个问题,在本文中,边缘属性明智的划分提出了R-GCN。所提出的分区划分模型和图形数据,以便R-GCN可以通过使用多个GPU来加速。此外,所提出的方法可以应用于单个GPU上的顺序执行。这两种情况都可以加速具有大型图形数据的R-GCN处理,其中原始模型无法在不分区的情况下放入单个GPU的设备内存中。实验结果表明,我们的分区方法使用四个GPU将R-GCN加速高达3.28倍,与CPU执行具有超过160万个节点和500万条边的数据集相比。此外,与具有80万个节点和200万条边的数据集的CPU执行相比,所提出的方法即使使用单个GPU也可以加速执行1.55倍。
R-GCN (Relational Graph Convolutional Network) is one of GNNs (Graph Neural Networks). The model tries predicting latent information by considering directions and types of edges in graph-structured data, such as knowledge bases. The model builds weight matrices to each edge attribute. Thus, the size of the neural network increases linearly with the number of edge types. Although GPUs can be used for accelerating the R-GCN processing, there is a possibility that the size of weight matrices exceeds GPU device memory. To address this issue, in this paper, an edge attribute-wise partitioning is proposed for R-GCN. The proposed partitioning divides the model and graph data so that R-GCN can be accelerated by using multiple GPUs. Also, the proposed approach can be applied to sequential execution on a single GPU. Both the cases can accelerate the R-GCN processing with large graph data, where the original model cannot be fit into a device memory of a single GPU without partitioning. Experimental results demonstrate that our partitioning method accelerates R-GCN by up to 3.28 times using four GPUs compared to CPU execution for a dataset with more than 1.6 million nodes and 5 million edges. Also, the proposed approach can accelerate the execution even with a single GPU by 1.55 times compared to the CPU execution for a dataset with 0.8 million nodes and 2 million edges.