A Semisupervised Classification Approach for Multidomain Networks with Domain Selection

A Semisupervised Classification Approach for Multidomain Networks with Domain Selection
复制标题

具有域选择的多域网络半监督分类方法

DOI:
10.1109/tnnls.2018.2837166
复制
发表时间:
2019
影响因子:
10.4
通讯作者:
Ng Michael K
Ng Michael K
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chen Chuan;Xin Jingxue;Wang Yong;Chen Luonan;Ng Michael K

文献摘要

被引文献

相似文献

多域网络分类在数据集成和机器学习中引起了极大的关注,它可以通过集成不同来源的信息来增强网络分类或预测性能。尽管之前取得了成功,但现有的多域网络学习方法通​​常假设同一组实例可以使用不同的视图,因此,它们寻求所有域的一致分类结果。然而,在许多现实问题中,每个域都有其特定的实例集,一个域中的一个实例可能对应于另一个域中的多个实例。此外,由于数据源的快速增长,不同的领域可能彼此不相关,这就要求选择与目标/聚焦领域相关的领域。这种设置下的一个关键挑战是如何在不丢失数据信息的情况下通过整合不同的数据表示来实现准确的预测。在本文中,我们提出了一种基于标签传播的多域网络半监督分类方法,即带有域选择的多域分类(MCS),它可以处理跨域信息和域中的不同实例集。特别是,凭借稀疏权重属性,所提出的 MCS 可以通过为它们分配比其他不相关域更高的权重来自动识别与我们的目标域相关的域。这不仅显着提高了分类精度,而且有助于获得目标域的最佳网络划分。从理论角度来看,我们等效地将MCS分解为两个具有解析解的更简单的子问题,可以通过它们的计算过程有效地求解。对合成数据集和真实世界数据集的大量实验结果从经验上证明了所提出的方法在预测性能和域选择能力方面的优势。
Multidomain network classification has attracted significant attention in data integration and machine learning, which can enhance network classification or prediction performance by integrating information from different sources. Despite the previous success, existing multidomain network learning methods usually assume that different views are available for the same set of instances, and thus, they seek a consistent classification result for all domains. However, in many real-world problems, each domain has its specific instance set, and one instance in one domain may correspond to multiple instances in another domain. Moreover, due to the rapid growth of data sources, different domains may not be relevant to each other, which asks for selecting domains relevant to the target/focused domain. A key challenge under this setting is how to achieve accurate prediction by integrating different data representations without losing data information. In this paper, we propose a semisupervised classification approach for a multidomain network based on label propagation, i.e., multidomain classification with domain selection (MCS), which can deal with the cross-domain information and different instance sets in domains. In particular, with sparse weight properties, the proposed MCS can automatically identify those domains relevant to our target domain by assigning them higher weights than the other irrelevant domains. This not only significantly improves a classification accuracy but also helps to obtain optimal network partition for the target domain. From the theoretical viewpoint, we equivalently decompose MCS into two simpler subproblems with analytical solutions, which can be efficiently solved by their computational procedures. Extensive experimental results on both synthetic and real-world data sets empirically demonstrate the advantages of the proposed approach in terms of both prediction performance and domain selection ability.