Co-Modality Graph Contrastive Learning for Imbalanced Node Classification

Co-Modality Graph Contrastive Learning for Imbalanced Node Classification
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Y. Qian;Chunhui Zhang;Yiming Zhang;Qianlong Wen;Yanfang Ye;Chuxu Zhang
Y. Qian;Chunhui Zhang;Yiming Zhang;Qianlong Wen;Yanfang Ye;Chuxu Zhang
中科院分区:
其他
文献类型:
--
作者:
Y. Qian;Chunhui Zhang;Yiming Zhang;Qianlong Wen;Yanfang Ye;Chuxu Zhang

文献摘要

被引文献

相似文献

图对比学习(GCL)利用图增强将图转换为不同的视图并进一步训练图神经网络(GNN),在图基准数据集上取得了相当大的成功。然而,将现有的 GCL 方法直接应用于现实世界数据仍然存在一些差距。首先,手工制作的图形增强需要反复试验,但仍然无法在多个任务上产生一致的性能。其次,大多数现实世界的图数据呈现类不平衡分布,但现有的 GCL 方法也不能免受数据不平衡的影响。因此,这项工作建议通过称为共模图对比学习(CM-GCL)的原则框架来明确应对这些挑战,以自动生成对比对并进一步学习未标记数据的平衡表示。具体来说,我们设计了模态间 GCL,以根据丰富的节点内容自动生成对比对(例如节点文本)。受到少数样本可以通过修剪深度神经网络而被“遗忘”这一事实的启发,我们自然地将网络修剪扩展到我们的 GCL 框架中以挖掘少数节点。基于此,我们通过将相应的节点-文本对推到一起,并将不相关的节点-文本对推开,以不同的方式共同训练两个修剪编码器(例如,GNN 和文本编码器)。同时,我们通过共同训练非剪枝 GNN 和剪枝 GNN 提出模态内 GCL,以确保具有相似属性特征的节点嵌入保持封闭。最后,我们在下游类不平衡节点分类任务上微调 GNN 编码器。大量的实验表明,我们的模型显着优于最先进的基线模型,并在现实世界的图表上学习更平衡的表示。我们的源代码可在 https://github.com/graphprojects/CM-GCL 获取。
Graph contrastive learning (GCL), leveraging graph augmentations to convert graphs into different views and further train graph neural networks (GNNs), has achieved considerable success on graph benchmark datasets. Yet, there are still some gaps in directly applying existing GCL methods to real-world data. First, handcrafted graph augmentations require trials and errors, but still can not yield consistent performance on multiple tasks. Second, most real-world graph data present class-imbalanced distribution but existing GCL methods are not immune to data imbalance. Therefore, this work proposes to explicitly tackle these challenges, via a principled framework called C o-M odality G raph C ontrastive L earning ( CM-GCL ) to automatically generate contrastive pairs and further learn balanced representation over unlabeled data. Specifically, we design inter-modality GCL to automatically generate contrastive pairs (e.g., node-text) based on rich node content. Inspired by the fact that minority samples can be “forgotten” by pruning deep neural networks, we naturally extend network pruning to our GCL framework for mining minority nodes. Based on this, we co-train two pruned encoders (e.g., GNN and text encoder) in different modalities by pushing the corresponding node-text pairs together and the irrelevant node-text pairs away. Meanwhile, we propose intra-modality GCL by co-training non-pruned GNN and pruned GNN, to ensure node embeddings with similar attribute features stay closed. Last, we fine-tune the GNN encoder on downstream class-imbalanced node classification tasks. Extensive experiments demonstrate that our model significantly outperforms state-of-the-art baseline models and learns more balanced representations on real-world graphs. Our source code is available at https://github.com/graphprojects/CM-GCL.