Theoretic Analysis and Extremely Easy Algorithms for Domain Adaptive Feature Learning

Theoretic Analysis and Extremely Easy Algorithms for Domain Adaptive Feature Learning
复制标题

DOI:
10.24963/ijcai.2017/272
复制
发表时间:
2015-09
期刊:
--
影响因子:
--
通讯作者:
Wenhao Jiang;Cheng Deng;W. Liu;F. Nie;K. F. Chung;Heng Huang
Wenhao Jiang;Cheng Deng;W. Liu;F. Nie;K. F. Chung;Heng Huang
中科院分区:
其他
文献类型:
--
作者:
Wenhao Jiang;Cheng Deng;W. Liu;F. Nie;K. F. Chung;Heng Huang

文献摘要

被引文献

相似文献

领域自适应问题出现在各种应用中,其中来自文本{源}域的训练数据集和来自文本{目标}域的测试数据集通常遵循不同的分布。设计有效的学习模型来解决这些问题的主要困难在于如何弥合来源分布和目标分布之间的差距。在本文中,我们提供了与线性分类器结合用于领域自适应的特征学习算法的全面分析。我们的分析表明,为了获得良好的自适应性能,源域分布和目标域分布的二阶矩应该是相似的。在新分析的基础上,提出了一种新的极易实现的领域自适应特征学习算法。此外,我们的算法通过利用多层进行了扩展,从而得到了一个深度线性模型。我们在Amazon评论数据集和ECML/PKDD 2006发现挑战中的垃圾数据集上评估了所提出的算法的有效性。
Domain adaptation problems arise in a variety of applications, where a training dataset from the \textit{source} domain and a test dataset from the \textit{target} domain typically follow different distributions. The primary difficulty in designing effective learning models to solve such problems lies in how to bridge the gap between the source and target distributions. In this paper, we provide comprehensive analysis of feature learning algorithms used in conjunction with linear classifiers for domain adaptation. Our analysis shows that in order to achieve good adaptation performance, the second moments of the source domain distribution and target domain distribution should be similar. Based on our new analysis, a novel extremely easy feature learning algorithm for domain adaptation is proposed. Furthermore, our algorithm is extended by leveraging multiple layers, leading to a deep linear model. We evaluate the effectiveness of the proposed algorithms in terms of domain adaptation tasks on the Amazon review dataset and the spam dataset from the ECML/PKDD 2006 discovery challenge.