Active Learning with Multi-Granular Graph Auto-Encoder

Active Learning with Multi-Granular Graph Auto-Encoder
复制标题

DOI:
10.1109/icdm50108.2020.00125
复制
发表时间:
2020-11
期刊:
2020 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Yi He;Xu Yuan;N. Tzeng;Xindong Wu
Yi He;Xu Yuan;N. Tzeng;Xindong Wu
中科院分区:
其他
文献类型:
--
作者:
Yi He;Xu Yuan;N. Tzeng;Xindong Wu

文献摘要

相似文献

网络数据的预测建模有许多现实世界的应用,例如社交网络中的欺诈检测、生物医学网络中的药物发现、引文网络中的论文主题分类等等。尽管先进的机器学习方法可以帮助建立相当准确的预测模型,但它们的适用性受到数据标记任务的极大阻碍,这些任务繁重、耗时且容易出错。在本文中,我们提出了一种新颖的网络数据主动学习范式,称为拓扑和内容感知(TACA)主动学习,旨在最大限度地减少标签数量,同时实现理想的模型精度水平。总体而言,TACA 从两个方面改进了现有工作:(1)TACA 对网络属性不做任何假设,而大多数现有工作仅在局部一致的网络上有效执行,其中链接的节点预计共享相同的标签;(2)TACA 生成查询而不依赖于模型性能,从而即使在查询的标签中存在噪声时也能享受鲁棒的预测结果。理论和经验证据均已提出,证实了我们方法的有效性和乐观性。
Predictive modeling of networked data finds many real-world applications, such as fraud detection in social networks, drug discovery in biomedical networks, paper topic classification in citation networks, and so forth. Although the advanced machine learning approaches can help build reasonably accurate predictive models, their applicability is immensely hindered by the data labeling tasks, which are onerous, time-consuming, and error-prone. In this paper, we propose a novel active learning paradigm for networked data, named topology-and-content-aware (TACA) active learning, aiming to minimize the number of labels while achieving a desirable level of model accuracy. Overall, TACA advances existing works from two aspects: (1) TACA makes no assumption on the network property, whereas most existing works only perform effectively on a locally consistent network in which linked nodes are expected to share the same labels and (2) TACA generates queries without relying on model performance, thereby enjoying robust predictive results even when noises exist in the queried labels. Both theoretical and empirical evidences are presented, substantiating the effectiveness of and optimism our approach.