Few-shot Node Classification with Extremely Weak Supervision

Few-shot Node Classification with Extremely Weak Supervision
复制标题

DOI:
10.1145/3539597.3570435
复制
发表时间:
2023-01
期刊:
Proceedings of the Sixteenth ACM International Conference on Web Search and Data Mining
影响因子:
--
通讯作者:
Song Wang;Yushun Dong;Kaize Ding;Chen Chen-Chen;Jundong Li
Song Wang;Yushun Dong;Kaize Ding;Chen Chen-Chen;Jundong Li
中科院分区:
其他
文献类型:
--
作者:
Song Wang;Yushun Dong;Kaize Ding;Chen Chen-Chen;Jundong Li

文献摘要

相似文献

Few-shot 节点分类旨在以有限的标记节点作为参考对节点进行分类。最近的少样本节点分类方法通常从具有丰富标记节点的类(即元训练类)中学习,然后推广到具有有限标记节点的类(即元测试类)。然而,在现实世界的图上,通常很难为许多类获得丰富的标记节点。在实践中,每个元训练类只能由几个标记节点组成,称为极弱监督问题。在少样本节点分类中,用于元训练的标记节点极其有限,元训练和元测试之间的泛化差距将变得更大,从而导致性能次优。为了解决这个问题,我们研究了一个监督极弱的少样本节点分类的新问题,并在流行的元学习框架下提出了一个原则框架 X-FNC。具体来说,我们的目标是在监督极弱的情况下积累不同元训练任务的元知识,并将这些知识推广到元测试任务。为了解决极度稀缺的标记节点带来的挑战,我们提出了两个基本模块来获取伪标记节点作为额外参考,并有效地从极其有限的监督信息中学习。我们进一步在监督极弱的四个节点分类数据集上进行了广泛的实验,以验证我们的框架与最先进的基线相比的优越性。
Few-shot node classification aims at classifying nodes with limited labeled nodes as references. Recent few-shot node classification methods typically learn from classes with abundant labeled nodes (i.e., meta-training classes) and then generalize to classes with limited labeled nodes (i.e., meta-test classes). Nevertheless, on real-world graphs, it is usually difficult to obtain abundant labeled nodes for many classes. In practice, each meta-training class can only consist of several labeled nodes, known as the extremely weak supervision problem. In few-shot node classification, with extremely limited labeled nodes for meta-training, the generalization gap between meta-training and meta-test will become larger and thus lead to suboptimal performance. To tackle this issue, we study a novel problem of few-shot node classification with extremely weak supervision and propose a principled framework X-FNC under the prevalent meta-learning framework. Specifically, our goal is to accumulate meta-knowledge across different meta-training tasks with extremely weak supervision and generalize such knowledge to meta-test tasks. To address the challenges resulting from extremely scarce labeled nodes, we propose two essential modules to obtain pseudo-labeled nodes as extra references and effectively learn from extremely limited supervision information. We further conduct extensive experiments on four node classification datasets with extremely weak supervision to validate the superiority of our framework compared to the state-of-the-art baselines.