OpenWGL: open-world graph learning for unseen class node classification

OpenWGL: open-world graph learning for unseen class node classification
复制标题

DOI:
10.1007/s10115-021-01594-0
复制
发表时间:
2021-08
影响因子:
2.7
通讯作者:
Man Wu;Shirui Pan;Xingquan Zhu
Man Wu;Shirui Pan;Xingquan Zhu
中科院分区:
计算机科学4区
文献类型:
--
作者:
Man Wu;Shirui Pan;Xingquan Zhu

文献摘要

相似文献

图学习,如节点分类,通常是在封闭世界环境中进行的。许多节点被标记,学习目标是正确地将剩余的(未标记的)节点分类为类,由标记的节点表示。在现实中,由于网络有限的标注能力或网络的动态演化特性,网络中的一些节点可能不属于任何现有的/见过的类,因此不能被封闭世界学习算法正确分类。在本文中,我们提出了一种新的开放世界图学习范式,其学习目标是将属于标记类的节点正确分类到正确的类别中,并将不属于标记类的节点分类到不可见的类中。开放世界图学习面临三大挑战:(1)图没有特征来表示学习的节点;(2)看不见的类节点没有标签,可能以不同于有标签的类的任意形式存在;(3)图学习应该区分一个节点是属于一个存在的/可见的类还是不可见的类。为了解决这些挑战,我们提出了一种不确定节点表示学习原理,使用多个版本的节点特征表示来测试分类器在节点上的响应,通过这种方法我们可以区分节点是否属于未知类。在技术方面,我们提出了约束变分图自编码器,使用标签损失和类不确定性损失约束,以确保节点表示学习对未知类敏感。因此,节点嵌入特征用分布表示,而不是确定性特征向量。为了检验节点属于可见类的确定性,提出了一种采样过程,生成多个版本的特征向量来表示每个节点,并使用自动阈值将不属于可见类的节点拒绝为未见类节点。通过在四个真实网络上使用图卷积网络和图关注网络的实验,验证了该算法的性能。实例研究和分析也表明了不确定表示学习和自动阈值选择在开放世界图学习中的优势。
Graph learning, such as node classification, is typically carried out in aclosed-worldsetting. A number of nodes are labeled, and the learning goal is to correctly classify remaining (unlabeled) nodes into classes, represented by the labeled nodes. In reality, due to limited labeling capability or dynamic evolving nature of networks, some nodes in the networks may not belong to any existing/seen classes and therefore cannot be correctly classified by closed-world learning algorithms. In this paper, we propose a newopen-worldgraph learning paradigm, where the learning goal is to correctly classify nodes belonging to labeled classes into correct categories and also classify nodes not belonging to labeled classes to an unseen class. Open-world graph learning has three major challenges: (1) Graphs do not have features to represent nodes for learning; (2) unseen class nodes do not have labels and may exist in an arbitrary form different from labeled classes; and (3) graph learning should differentiate whether a node belongs to an existing/seen class or an unseen class. To tackle the challenges, we propose an uncertain node representation learning principle to use multiple versions of node feature representation to test a classifier’s response on a node, through which we can differentiate whether a node belongs to the unseen class. Technical wise, we propose constrained variational graph autoencoder, using label loss and class uncertainty loss constraints, to ensure that node representation learning is sensitive to the unseen class. As a result, node embedding features are denoted by distributions, instead of deterministic feature vectors. In order to test the certainty of a node belonging to seen classes, a sampling process is proposed to generate multiple versions of feature vectors to represent each node, using automatic thresholding to reject nodes not belonging to seen classes as unseen class nodes. Experiments, using graph convolutional networks and graph attention networks on four real-world networks, demonstrate the algorithm performance. Case studies and ablation analysis also show the advantage of the uncertain representation learning and automatic threshold selection for open-world graph learning.