Online adaptation to post-edits for phrase-based statistical machine translation

Online adaptation to post-edits for phrase-based statistical machine translation
复制标题

DOI:
10.1007/s10590-014-9159-7
复制
发表时间:
2014-12
影响因子:
1.9
通讯作者:
N. Bertoldi;P. Simianer;M. Cettolo;K. Wäschle;Marcello Federico;S. Riezler
N. Bertoldi;P. Simianer;M. Cettolo;K. Wäschle;Marcello Federico;S. Riezler
中科院分区:
--
文献类型:
--
作者:
N. Bertoldi;P. Simianer;M. Cettolo;K. Wäschle;Marcello Federico;S. Riezler

文献摘要

被引文献

相似文献

最近的研究表明,人工翻译的准确性和速度可以从机器翻译系统的编辑后输出中受益,更高质量的输出收益更大。我们提出了一个有效的在线学习框架,使基于短语的统计机器翻译系统的所有模块适应后编辑的翻译。我们使用约束搜索技术从后期编辑中提取新的短语翻译,而不需要重新对齐,并且在不需要替代引用的情况下提取用于判别训练的短语对特征。此外,基于从后期编辑中提取的图构建了基于缓存的语言模型。我们在模拟的后期编辑场景和现场测试数据中提出了实验结果。每个单独的模块都大大提高了翻译质量。这些模块可以有效地实现,并允许直接堆叠,在几个翻译方向和领域上产生显着的加法改进。
Recent research has shown that accuracy and speed of human translators can benefit frompost-editingoutput of machine translation systems, with larger benefits for higher quality output. We present an efficient online learning framework for adapting all modules of a phrase-based statistical machine translation system to post-edited translations. We use a constrained search technique to extract new phrase-translations from post-edits without the need of re-alignments, and to extract phrase pair features for discriminative training without the need for surrogate references. In addition, a cache-based language model is built on-grams extracted from post-edits. We present experimental results in a simulated post-editing scenario and on field-test data. Each individual module substantially improves translation quality. The modules can be implemented efficiently and allow for a straightforward stacking, yielding significant additive improvements on several translation directions and domains.