Cross-Domain Learning from Multiple Sources: A Consensus Regularization Perspective

Cross-Domain Learning from Multiple Sources: A Consensus Regularization Perspective
复制标题

DOI:
10.1109/tkde.2009.205
复制
发表时间:
2010-12
影响因子:
8.9
通讯作者:
Fuzhen Zhuang;Ping Luo;Hui Xiong;Yuhong Xiong;Qing He;Zhongzhi Shi
Fuzhen Zhuang;Ping Luo;Hui Xiong;Yuhong Xiong;Qing He;Zhongzhi Shi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Fuzhen Zhuang;Ping Luo;Hui Xiong;Yuhong Xiong;Qing He;Zhongzhi Shi

文献摘要

被引文献

相似文献

跨域分类研究如何将学习模型从一个域调整到共享相似数据特征的另一个域。虽然有许多现有的作品沿着这条线,他们中的许多人只专注于学习从一个单一的源域到一个目标域。特别是,剩下的挑战是如何将从多个源域学到的知识应用到目标域。事实上,来自多个源域的数据可以在语义上相关,但具有不同的数据分布。目前还不清楚如何利用多个源域之间的分布差异,以提高在目标域的学习性能。为此,在本文中,我们提出了一个共识正则化框架,用于从多个源域到目标域的学习。在这个框架中,一个本地分类器的训练,同时考虑本地数据在一个源域和预测共识与从其他源域学习的分类器。此外,我们提供了一个理论分析,以及建议的共识正则化框架的实证研究。文本分类和图像分类的实验结果表明了该一致性正则化学习方法的有效性。最后,为了处理多个源域在地理上分布的情况,我们还开发了所提出的算法的分布式版本,这避免了需要将所有数据上传到一个集中的位置,并有助于减轻隐私问题。
Classification across different domains studies how to adapt a learning model from one domain to another domain which shares similar data characteristics. While there are a number of existing works along this line, many of them are only focused on learning from a single source domain to a target domain. In particular, a remaining challenge is how to apply the knowledge learned from multiple source domains to a target domain. Indeed, data from multiple source domains can be semantically related, but have different data distributions. It is not clear how to exploit the distribution differences among multiple source domains to boost the learning performance in a target domain. To that end, in this paper, we propose a consensus regularization framework for learning from multiple source domains to a target domain. In this framework, a local classifier is trained by considering both local data available in one source domain and the prediction consensus with the classifiers learned from other source domains. Moreover, we provide a theoretical analysis as well as an empirical study of the proposed consensus regularization framework. The experimental results on text categorization and image classification problems show the effectiveness of this consensus regularization learning method. Finally, to deal with the situation that the multiple source domains are geographically distributed, we also develop the distributed version of the proposed algorithm, which avoids the need to upload all the data to a centralized location and helps to mitigate privacy concerns.