Integrating an Unsupervised Transliteration Model into Statistical Machine Translation

Integrating an Unsupervised Transliteration Model into Statistical Machine Translation
复制标题

将无监督音译模型集成到统计机器翻译中

DOI:
--
复制
发表时间:
2014
期刊:
Conference of the European Chapter of the Association for Computational Linguistics
影响因子:
--
通讯作者:
Philipp Koehn
Philipp Koehn
中科院分区:
--
文献类型:
--
作者:
Nadir Durrani;Hassan Sajjad;Hieu D. Hoang;Philipp Koehn

文献摘要

被引文献

相似文献

我们研究了三种将无监督音译模型集成到端到端SMT系统中的方法。我们从平行数据中归纳出一个音译模型,并将其用于OOV词的翻译。我们的方法是完全无人监督的,并且独立于语言。在整合音译的方法中,我们观察到7种语言对的BLEU分数从0.23-0.75(0.41)提高。我们还表明,与黄金标准音译语料库相比,我们挖掘的音译语料库提供了更好的规则覆盖率和翻译质量。
We investigate three methods for integrating an unsupervised transliteration model into an end-to-end SMT system. We induce a transliteration model from parallel data and use it to translate OOV words. Our approach is fully unsupervised and language independent. In the methods to integrate transliterations, we observed improvements from 0.23-0.75 ( 0.41) BLEU points across 7 language pairs. We also show that our mined transliteration corpora provide better rule coverage and translation quality compared to the gold standard transliteration corpora.