Improving Statistical Machine Translation Using Word Sense Disambiguation

Improving Statistical Machine Translation Using Word Sense Disambiguation
复制标题

DOI:
--
复制
发表时间:
2007-06
期刊:
--
影响因子:
--
通讯作者:
Marine Carpuat;Dekai Wu
Marine Carpuat;Dekai Wu
中科院分区:
其他
文献类型:
--
作者:
Marine Carpuat;Dekai Wu

文献摘要

被引文献

相似文献

我们首次证明,将词义消歧系统的预测纳入典型的基于短语的统计机器翻译(SMT)模型中,可以持续提高所有三种不同的IWTECH汉英测试集的翻译质量,并在更大的NIST汉英MT任务中产生统计上的显着改善-而且不会损害任何测试集的性能,不仅根据BLEU,而且根据所有八个最常用的自动评估指标。最近的工作已经挑战的假设,词义消歧(WSD)系统是有用的SMT。然而,词汇选择不准确仍是影响SMT翻译质量的主要因素。在本文中,我们解决这个问题,通过调查一个新的策略,将WSD整合到SMT系统,进行充分的短语多词消歧。而不是直接将一个感官风格的WSD系统,我们重新定义的WSD任务,以匹配完全相同的短语翻译消歧任务所面临的短语为基础的SMT系统。我们的研究结果提供了第一个已知的经验证据,词汇语义确实是有用的SMT,尽管声称相反。
We show for the first time that incorporating the predictions of a word sense disambiguation system within a typical phrase-based statistical machine translation (SMT) model consistently improves translation quality across all three different IWSLT ChineseEnglish test sets, as well as producing statistically significant improvements on the larger NIST Chinese-English MT task— and moreover never hurts performance on any test set, according not only to BLEU but to all eight most commonly used automatic evaluation metrics. Recent work has challenged the assumption that word sense disambiguation (WSD) systems are useful for SMT. Yet SMT translation quality still obviously suffers from inaccurate lexical choice. In this paper, we address this problem by investigating a new strategy for integrating WSD into an SMT system, that performs fully phrasal multi-word disambiguation. Instead of directly incorporating a Senseval-style WSD system, we redefine the WSD task to match the exact same phrasal translation disambiguation task faced by phrase-based SMT systems. Our results provide the first known empirical evidence that lexical semantics are indeed useful for SMT, despite claims to the contrary.