Multi-domain learning by confidence-weighted parameter combination

Multi-domain learning by confidence-weighted parameter combination
复制标题

DOI:
10.1007/s10994-009-5148-0
复制
发表时间:
2010-05-01
期刊:
影响因子:
7.5
通讯作者:
Crammer, Koby
Crammer, Koby
中科院分区:
计算机科学3区
文献类型:
--
作者:
Dredze, Mark;Kulesza, Alex;Crammer, Koby

文献摘要

被引文献

相似文献

用于各种任务的最先进的统计NLP系统通常从特定领域的标记训练数据中学习。然而,可能有多个领域或兴趣源,系统必须执行这些领域或兴趣源。例如,垃圾邮件过滤系统必须为许多用户提供高质量的预测,每个用户接收来自不同来源的电子邮件,并且可能对哪些是垃圾邮件或哪些不是垃圾邮件做出略微不同的决定。我们不是为每个领域学习单独的模型,而是探索跨多个领域学习的系统。提出了一种基于多分类器参数组合的多领域在线学习框架。我们的算法借鉴了多任务学习和领域自适应,使多个源领域分类器适应新的目标领域,跨多个相似领域学习,以及跨大量不同领域学习。我们在两个流行的NLP领域自适应任务上评估了我们的算法:情感分类和垃圾邮件过滤。
State-of-the-art statistical NLP systems for a variety of tasks learn from labeled training data that is often domain specific. However, there may be multiple domains or sources of interest on which the system must perform. For example, a spam filtering system must give high quality predictions for many users, each of whom receives emails from different sources and may make slightly different decisions about what is or is not spam. Rather than learning separate models for each domain, we explore systems that learn across multiple domains. We develop a new multi-domain online learning framework based on parameter combination from multiple classifiers. Our algorithms draw from multi-task learning and domain adaptation to adapt multiple source domain classifiers to a new target domain, learn across multiple similar domains, and learn across a large number of disparate domains. We evaluate our algorithms on two popular NLP domain adaptation tasks: sentiment classification and spam filtering.