Domain adaptation for statistical classifiers

Domain adaptation for statistical classifiers
复制标题

DOI:
10.1613/jair.1872
复制
发表时间:
2006-01-01
影响因子:
5
通讯作者:
Marcu, D
Marcu, D
中科院分区:
计算机科学3区
文献类型:
--
作者:
Daumé, H;Marcu, D

文献摘要

被引文献

相似文献

统计学习理论中使用的最基本假设是训练数据和测试数据来自相同的底层分布。不幸的是,在许多应用程序中,“域内”测试数据是从与训练数据的“域外”分布相关但不相同的分布中提取的。我们考虑了标记域外数据丰富而标记域内数据稀缺的常见情况。我们在一个简单的混合模型中引入了这个问题的统计公式,并给出了这个框架的一个实例,以最大熵分类器和它们的线性链对应。基于条件期望最大化技术,提出了这种特殊情况下的有效推理算法。我们的实验结果表明,我们的方法可以在自然语言处理领域的四个不同数据集上的三个真实世界任务上提高性能。
The most basic assumption used in statistical learning theory is that training data and test data are drawn from the same underlying distribution. Unfortunately, in many applications, the "in-domain" test data is drawn from a distribution that is related, but not identical, to the "out-of-domain" distribution of the training data. We consider the common case in which labeled out-of-domain data is plentiful, but labeled in-domain data is scarce. We introduce a statistical formulation of this problem in terms of a simple mixture model and present an instantiation of this framework to maximum entropy classifiers and their linear chain counterparts. We present efficient inference algorithms for this special case based on the technique of conditional expectation maximization. Our experimental results show that our approach leads to improved performance on three real world tasks on four different data sets from the natural language processing domain.