Collaborative Dual-PLSA: mining distinction and commonality across multiple domains for text classification

Collaborative Dual-PLSA: mining distinction and commonality across multiple domains for text classification
复制标题

DOI:
10.1145/1871437.1871486
复制
发表时间:
2010-10
期刊:
Proceedings of the 19th ACM international conference on Information and knowledge management
影响因子:
--
通讯作者:
Fuzhen Zhuang;Ping Luo;Zhiyong Shen;Qing He;Yuhong Xiong;Zhongzhi Shi;Hui Xiong
Fuzhen Zhuang;Ping Luo;Zhiyong Shen;Qing He;Yuhong Xiong;Zhongzhi Shi;Hui Xiong
中科院分区:
其他
文献类型:
--
作者:
Fuzhen Zhuang;Ping Luo;Zhiyong Shen;Qing He;Yuhong Xiong;Zhongzhi Shi;Hui Xiong

文献摘要

被引文献

相似文献

跨域文本分类问题考虑了多个数据域之间的分布差异。在这项研究中,我们沿着这条线展示了两个新的观察结果。首先,数据分布的差异可能来自不同领域使用不同的关键词来表达相同的概念。第二,这个概念特征和文档类之间的关联可以跨域稳定。这两个问题实际上是跨数据域的区别和共性。受上述观察结果的启发,我们提出了一个生成统计模型,名为协作双PLSA(CD-PLSA),同时捕获多个域之间的域区别和共性。与概率潜在语义分析(PLSA)只有一个潜在变量不同,该模型有两个潜在因子y和z,分别对应于词的概念和文档类别。共享的共性与多个领域的差异交织在一起,也被用作知识转化的桥梁。我们利用期望最大化(EM)算法来学习这个模型,并提出其分布式版本来处理数据域在地理上相互分离的情况。最后,我们对数百个具有多个源域和多个目标域的分类任务进行了广泛的实验,以验证所提出的CD-PLSA模型优于现有的最先进的监督和迁移学习方法。特别是,我们表明,CD-PLSA是更宽容的分布差异。
The distribution difference among multiple data domains has been considered for the cross-domain text classification problem. In this study, we show two new observations along this line. First, the data distribution difference may come from the fact that different domains use different key words to express the same concept. Second, the association between this conceptual feature and the document class may be stable across domains. These two issues are actually the distinction and commonality across data domains. Inspired by the above observations, we propose a generative statistical model, named Collaborative Dual-PLSA (CD-PLSA), to simultaneously capture both the domain distinction and commonality among multiple domains. Different from Probabilistic Latent Semantic Analysis (PLSA) with only one latent variable, the proposed model has two latent factors y and z, corresponding to word concept and document class respectively. The shared commonality intertwines with the distinctions over multiple domains, and is also used as the bridge for knowledge transformation. We exploit an Expectation Maximization (EM) algorithm to learn this model, and also propose its distributed version to handle the situation where the data domains are geographically separated from each other. Finally, we conduct extensive experiments over hundreds of classification tasks with multiple source domains and multiple target domains to validate the superiority of the proposed CD-PLSA model over existing state-of-the-art methods of supervised and transfer learning. In particular, we show that CD-PLSA is more tolerant of distribution differences.