Robust Network Enhancement From Flawed Networks

Robust Network Enhancement From Flawed Networks
复制标题

DOI:
10.1109/tkde.2020.3025147
复制
发表时间:
2022-07
影响因子:
8.9
通讯作者:
Jiarong Xu;Yang Yang-Yang;Chunping Wang;Zongtao Liu;Jing Zhang;Lei Chen;Jiangang Lu
Jiarong Xu;Yang Yang-Yang;Chunping Wang;Zongtao Liu;Jing Zhang;Lei Chen;Jiangang Lu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jiarong Xu;Yang Yang-Yang;Chunping Wang;Zongtao Liu;Jing Zhang;Lei Chen;Jiangang Lu

文献摘要

相似文献

现实世界中的网络数据由于采样不完全、测量不完善等原因,往往容易出错;当对这些有缺陷的网络执行网络分析或建模(例如节点分类和链路预测)时,这又导致不准确的结果。在本文中,我们的目标是重建一个可靠的网络从一个有缺陷的,无向的,未加权的网络,一个过程称为网络增强。更具体地说,网络增强旨在检测在网络中观察到但不应该存在于真实的世界中的噪声链路,以及预测在真实的世界中确实存在但仍未观察到的缺失链路。虽然已经进行了一些尝试来检测噪声链接或丢失的链接,但很少有这些作品考虑统一这两个任务,即使它们是相互依赖的,并且能够相互促进彼此的性能。因此,在本文中,我们提出了E-Net,一个端到端的图神经网络模型,利用这两个任务的相互影响,以更有效地实现这两个目标。一方面,检测噪声链路可以有益于丢失链路预测的性能,而另一方面,预测丢失链路可以在这些噪声链路的标签不可用时为检测噪声链路检测提供间接监督。此外,通过提出一种基于重启随机游走的子图提取机制,该模型可以扩展到大型网络,并且能够保持局部和全局结构特征。在几种类型的大型网络上的实验结果表明,所提出的模型获得了10.7%的改善,平均在F1的预测丢失的链接,沿着与平均3.7%的改善,在检测噪声链接的精度相比,国家的最先进的基线。
Network data in real-world tends to be error-prone due to incomplete sampling, imperfect measurements, etc.; this in turn results in inaccurate results when performing network analysis or modeling, such as node classification and link prediction, on these flawed networks. In this paper, we aim to reconstruct a reliable network from a flawed, undirected, unweighted network, a process referred to network enhancement. More specifically, network enhancement aims to detect the noisy links that are observed in the network but should not exist in the real world, as well as to predict the missing links that do indeed exist in the real world yet remain unobserved. While some attempts have been made to detect either noisy links or missing links, few of these works have considered unifying these two tasks, even though they are inter-dependent and capable of mutually boosting each others’ performance. In this paper, we therefore propose E-Net, an end-to-end graph neural network model, to leverage the mutual influence of these two tasks in order to achieve both goals more effectively. On one hand, detecting noisy links can benefit the performance of missing link prediction, while on the other hand, predicting missing links can provide indirect supervision for detecting noisy link detection when the labels of these noisy links are unavailable. Moreover, by proposing a subgraph extraction mechanism based on random walk with restart, the model can be scaled up to large networks and is able to preserve the local and global structural characteristics. The experimental results on several types of large networks demonstrate that the proposed model obtains an improvement of 10.7 percent on average in terms of F1 for predicting missing links, along with an average of 3.7 percent improvement in terms of precision for detecting noisy links compared with the state-of-the-art baselines.