Investigations on Translation Model Adaptation Using Monolingual Data

Investigations on Translation Model Adaptation Using Monolingual Data
复制标题

DOI:
--
复制
发表时间:
2011-07
期刊:
--
影响因子:
--
通讯作者:
Patrik Lambert;Holger Schwenk;Christophe Servan;Sadaf Abdul Rauf
Patrik Lambert;Holger Schwenk;Christophe Servan;Sadaf Abdul Rauf
中科院分区:
其他
文献类型:
--
作者:
Patrik Lambert;Holger Schwenk;Christophe Servan;Sadaf Abdul Rauf

文献摘要

被引文献

相似文献

大多数用于训练统计机器翻译系统的翻译模型的免费并行数据来自非常特定的来源(欧洲议会,联合国等)。因此,人们对执行翻译模型的适配的方法越来越感兴趣。一种流行的方法是基于无监督训练,也称为自我增强。两者都只使用单语数据来适应翻译模型。在本文中,我们扩展了以前的工作,并提供了新的见解,在现有的方法。我们报告结果的法语和英语之间的翻译。对于在超过2.8亿字的人类翻译并行数据上训练的非常有竞争力的基线,观察到高达0.5 BLEU的改进。
Most of the freely available parallel data to train the translation model of a statistical machine translation system comes from very specific sources (European parliament, United Nations, etc). Therefore, there is increasing interest in methods to perform an adaptation of the translation model. A popular approach is based on unsupervised training, also called self-enhancing. Both only use monolingual data to adapt the translation model. In this paper we extend the previous work and provide new insight in the existing methods. We report results on the translation between French and English. Improvements of up to 0.5 BLEU were observed with respect to a very competitive baseline trained on more than 280M words of human translated parallel data.