Active Learning for Node Classification using a Convex Optimization approach

Active Learning for Node Classification using a Convex Optimization approach
复制标题

DOI:
10.1109/bigdataservice55688.2022.00022
复制
发表时间:
2022-08
期刊:
2022 IEEE Eighth International Conference on Big Data Computing Service and Applications (BigDataService)
影响因子:
--
通讯作者:
D. Agarwal;Balasubramaniam Natarajan
D. Agarwal;Balasubramaniam Natarajan
中科院分区:
其他
文献类型:
--
作者:
D. Agarwal;Balasubramaniam Natarajan

文献摘要

相似文献

在工业4.0时代,与大数据分析相关的最新进展得益于基于神经网络(NN)架构的决策模型的开发和部署。除了在欧几里得空间中表示的数据,如图像,文本或视频,还有越来越多的应用程序需要在非欧几里得域中表示数据。这些数据通常表示为具有各种实体之间的复杂交互和相互依赖性的图。图数据的复杂性对传统的深度神经网络模型提出了巨大的挑战,图神经网络(GNN)是网络数据建模和分析的强大扩展。这些计算模型的训练需要大量的标记数据。主动学习(AL)通过在训练过程中选择信息量最大的实例进行标记来帮助克服这个问题。本文将AL与GNN相结合,对属性图中的节点进行半监督分类。AL框架被描绘成一个凸优化问题,采用基于差异的稀疏建模代表性选择(DSMRS)。实验评估表明,使用选定的图形特定的指标(中心性和鲁棒性措施)作为AL算法,导致分类性能提高高达10%。
The recent advancements related to big data analytics in the era of Industry 4.0 are fueled by development and deployment of decision models based on neural network (NN) architectures. In addition to the data represented in Euclidean space, like, images, text, or videos, there are increasing applications which demand data representation in non-Euclidean domains. Such data are typically represented as graphs with complex interactions and interdependencies between various entities. The complexity of graph data imposes substantial challenge on the traditional Deep NN models, and Graph Neural Networks (GNN) are a powerful extension for modeling and analysis of networked data. The training of these computational models requires large amounts of labeled data. Active Learning (AL) helps to overcome this issue by selecting the most informative instances for labeling during the training process. This paper combines AL with GNN for semi-supervised classification of nodes in attributed graphs. The AL framework is portrayed as a convex optimization problem by employing Dissimilarity-based Sparse Modeling Representative Selection (DSMRS). The experimental evaluation demonstrates that using the selected graph-specific metrics (centrality and robustness measures) as AL heuristics leads to an improvement in classification performance by upto 10%.