Feature selection: Key to enhance node classification with graph neural networks

Feature selection: Key to enhance node classification with graph neural networks
复制标题

特征选择:利用图神经网络增强节点分类的关键

DOI:
10.1049/cit2.12166
复制
发表时间:
2023
影响因子:
5.1
通讯作者:
Murata Tsuyoshi
Murata Tsuyoshi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Maurya Sunil Kumar;Liu Xin;Murata Tsuyoshi

文献摘要

相似文献

图表有助于定义数据中实体之间的关系。这些由边表示的关系通常提供额外的上下文信息,这些信息可用于发现数据中的模式。图神经网络(GNN)利用图结构的归纳偏差来学习和预测各种任务。图神经网络的主要操作是基于图的结构在节点的邻居上执行的特征聚合步骤。除了它自己的特征之外,对于每一跳,节点从它的邻居那里获得额外的组合特征。这些聚合的特征有助于定义节点相对于标签的相似性或不相似性,并且对于节点分类等任务很有用。然而,在真实的世界数据中,不同跳处的邻居的特征可能与节点的特征不相关。因此,GNN的任何不加选择的特征聚合都可能导致添加噪声特征,从而导致模型性能下降。在这项工作中,我们表明,选择性聚合的节点功能,从不同的跳导致更好的性能比默认聚合的节点分类任务。此外,我们提出了一个具有分类器模型和选择器模型的双网GNN架构。分类器模型在输入节点特征的子集上训练以预测节点标签,而选择器模型学习向分类器提供最佳输入子集以获得最佳性能。这两个模型被联合训练,以学习在节点标签预测中提供更高准确性的最佳特征子集。通过大量的实验,我们证明了我们提出的模型优于特征选择方法和最先进的GNN模型,显著提高了27.8%。
Graphs help to define the relationships between entities in the data. These relationships, represented by edges, often provide additional context information which can be utilised to discover patterns in the data. Graph Neural Networks (GNNs) employ the inductive bias of the graph structure to learn and predict on various tasks. The primary operation of graph neural networks is the feature aggregation step performed over neighbours of the node based on the structure of the graph. In addition to its own features, for each hop, the node gets additional combined features from its neighbours. These aggregated features help define the similarity or dissimilarity of the nodes with respect to the labels and are useful for tasks like node classification. However, in real‐world data, features of neighbours at different hops may not correlate with the node's features. Thus, any indiscriminate feature aggregation by GNN might cause the addition of noisy features leading to degradation in model's performance. In this work, we show that selective aggregation of node features from various hops leads to better performance than default aggregation on the node classification task. Furthermore, we propose a Dual‐Net GNN architecture with a classifier model and a selector model. The classifier model trains over a subset of input node features to predict node labels while the selector model learns to provide optimal input subset to the classifier for the best performance. These two models are trained jointly to learn the best subset of features that give higher accuracy in node label predictions. With extensive experiments, we show that our proposed model outperforms both feature selection methods and state‐of‐the‐art GNN models with remarkable improvements up to 27.8%.