Link Prediction with Non-Contrastive Learning

Link Prediction with Non-Contrastive Learning
复制标题

DOI:
10.48550/arxiv.2211.14394
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
William Shiao;Zhichun Guo;Tong Zhao;E. Papalexakis;Yozen Liu;Neil Shah
William Shiao;Zhichun Guo;Tong Zhao;E. Papalexakis;Yozen Liu;Neil Shah
中科院分区:
其他
文献类型:
--
作者:
William Shiao;Zhichun Guo;Tong Zhao;E. Papalexakis;Yozen Liu;Neil Shah

文献摘要

相似文献

图神经网络(GNN)空间中最近的一个焦点领域是图自监督学习(SSL),其目的是在没有标记数据的情况下导出有用的节点表示。值得注意的是,许多最先进的图SSL方法是对比方法,其使用正样本和负样本的组合来学习节点表示。由于负采样的挑战(缓慢和模型敏感性),最近的文献介绍了非对比方法,而不是只使用正样本。虽然这些方法在节点级任务中表现出很好的性能,但它们对链接预测任务的适用性尚未得到探索,这些任务涉及预测节点对之间的链接存在(并且对推荐系统上下文具有广泛的适用性)。在这项工作中,我们广泛地评估现有的非对比性方法的性能,在两个转导和感应设置的链接预测。虽然大多数现有的非对比方法整体表现不佳,但我们发现,令人惊讶的是,BGRL通常在转导设置中表现良好。然而,它在更现实的归纳设置中表现不佳,在这种设置中,模型必须泛化到指向/来自看不见的节点的链接。我们发现非对比模型倾向于过拟合训练图,并使用此分析提出T-BGRL,这是一种新的非对比框架,它包含廉价的腐败以提高模型的泛化能力。这个简单的修改极大地提高了我们5/6数据集的归纳性能,在50次点击中提高了120%--所有这些都与其他非对比基线相当,比性能最好的对比基线快14倍。我们的工作为链接预测的非对比学习提供了有趣的发现,并为未来的研究人员进一步扩展这一领域铺平了道路。
A recent focal area in the space of graph neural networks (GNNs) is graph self-supervised learning (SSL), which aims to derive useful node representations without labeled data. Notably, many state-of-the-art graph SSL methods are contrastive methods, which use a combination of positive and negative samples to learn node representations. Owing to challenges in negative sampling (slowness and model sensitivity), recent literature introduced non-contrastive methods, which instead only use positive samples. Though such methods have shown promising performance in node-level tasks, their suitability for link prediction tasks, which are concerned with predicting link existence between pairs of nodes (and have broad applicability to recommendation systems contexts) is yet unexplored. In this work, we extensively evaluate the performance of existing non-contrastive methods for link prediction in both transductive and inductive settings. While most existing non-contrastive methods perform poorly overall, we find that, surprisingly, BGRL generally performs well in transductive settings. However, it performs poorly in the more realistic inductive settings where the model has to generalize to links to/from unseen nodes. We find that non-contrastive models tend to overfit to the training graph and use this analysis to propose T-BGRL, a novel non-contrastive framework that incorporates cheap corruptions to improve the generalization ability of the model. This simple modification strongly improves inductive performance in 5/6 of our datasets, with up to a 120% improvement in Hits@50--all with comparable speed to other non-contrastive baselines and up to 14x faster than the best-performing contrastive baseline. Our work imparts interesting findings about non-contrastive learning for link prediction and paves the way for future researchers to further expand upon this area.