Rethinking and Scaling Up Graph Contrastive Learning: An Extremely Efficient Approach with Group Discrimination

Rethinking and Scaling Up Graph Contrastive Learning: An Extremely Efficient Approach with Group Discrimination
复制标题

DOI:
10.48550/arxiv.2206.01535
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yizhen Zheng;Shirui Pan;Vincent C. S. Lee;Yu Zheng;Philip S. Yu
Yizhen Zheng;Shirui Pan;Vincent C. S. Lee;Yu Zheng;Philip S. Yu
中科院分区:
其他
文献类型:
--
作者:
Yizhen Zheng;Shirui Pan;Vincent C. S. Lee;Yu Zheng;Philip S. Yu

文献摘要

被引文献

相似文献

图对比学习(GCL)通过自监督学习方案消除了图表示学习(GRL)对标签信息的严重依赖。其核心思想是通过最大化相似实例的互信息来学习,这需要计算两个节点实例之间的相似性。然而,GCL在时间和内存消耗方面效率低下。此外,GCL通常需要大量的训练epoch才能在大规模数据集上得到良好的训练。受技术缺陷观察的启发(即,针对GCL的两个代表性著作DGI和MVGRL中常用的Sigmoid函数使用不当的问题,本文重新审视了GCL,引入了一种新的自监督图表示学习范式,即群判别(GD),并提出了一种新的基于GD的图群判别(GGD)方法。代替相似度计算,GGD直接区分两组节点样本与一个非常简单的二进制交叉熵损失。此外,与GCL方法相比,GGD在大规模数据集上需要更少的训练时间来获得有竞争力的性能。这两个优点赋予GGD非常高效的性能。大量的实验表明,GGD在八个数据集上的性能优于最先进的自监督方法。特别是,GGD可以在ogbn-arxiv上在0.18秒(包括数据预处理在内为6.44秒)内完成训练,这比GCL基线快了几个数量级(10,000+),同时消耗的内存要少得多。经过9小时的ogbn-papers 100 M十亿边缘训练,GGD在准确性和效率方面都优于GCL同行。
Graph contrastive learning (GCL) alleviates the heavy reliance on label information for graph representation learning (GRL) via self-supervised learning schemes. The core idea is to learn by maximising mutual information for similar instances, which requires similarity computation between two node instances. However, GCL is inefficient in both time and memory consumption. In addition, GCL normally requires a large number of training epochs to be well-trained on large-scale datasets. Inspired by an observation of a technical defect (i.e., inappropriate usage of Sigmoid function) commonly used in two representative GCL works, DGI and MVGRL, we revisit GCL and introduce a new learning paradigm for self-supervised graph representation learning, namely, Group Discrimination (GD), and propose a novel GD-based method called Graph Group Discrimination (GGD). Instead of similarity computation, GGD directly discriminates two groups of node samples with a very simple binary cross-entropy loss. In addition, GGD requires much fewer training epochs to obtain competitive performance compared with GCL methods on large-scale datasets. These two advantages endow GGD with very efficient property. Extensive experiments show that GGD outperforms state-of-the-art self-supervised methods on eight datasets. In particular, GGD can be trained in 0.18 seconds (6.44 seconds including data preprocessing) on ogbn-arxiv, which is orders of magnitude (10,000+) faster than GCL baselines while consuming much less memory. Trained with 9 hours on ogbn-papers100M with billion edges, GGD outperforms its GCL counterparts in both accuracy and efficiency.