A Comprehensive Analysis of Bilingual Lexicon Induction

A Comprehensive Analysis of Bilingual Lexicon Induction
复制标题

DOI:
10.1162/coli_a_00284
复制
发表时间:
2017-06-01
影响因子:
9.3
通讯作者:
Callison-Burch, Chris
Callison-Burch, Chris
中科院分区:
计算机科学3区
文献类型:
--
作者:
Irvine, Ann;Callison-Burch, Chris

文献摘要

被引文献

相似文献

双语词汇归纳是指从单语语料库中归纳出两种语言的词汇翻译。在这篇文章中,我们提出了迄今为止最全面的双语词汇归纳分析。我们提出了各种语言和数据大小的实验。我们研究从25种外语翻译成英语:阿尔巴尼亚语,印地语,孟加拉语,波斯尼亚语,保加利亚语,宿务语,古吉拉特语,印地语,匈牙利语,印度尼西亚语,拉脱维亚语,尼泊尔语,罗马尼亚语,塞尔维亚语,斯洛伐克语,索马里语,西班牙语,瑞典语,泰米尔语,泰卢固语,土耳其语,乌克兰语,乌兹别克语,越南语和威尔士语。我们分析的行为双语词汇归纳的低频词,而不是只测试高频词,如以前的研究所做的。低频词与统计机器翻译更相关,其中系统通常缺乏对其训练数据之外的稀有词的翻译。我们系统地探讨了双语词汇归纳法发现的影响翻译质量的一系列特征和现象。我们提供了说明性的例子,排名最高的翻译正交信号的翻译等价,如上下文相似性和时间相似性。我们分析了频率和突发性,以及种子双语词典和单语训练语料库的大小的影响。此外,我们介绍了一种新的歧视性的方法,双语词汇归纳。我们的判别模型是能够结合各种各样的功能,单独提供翻译等价性的弱指标。当区别性地设置特征权重时,这些信号产生比以无监督方式组合信号的先前方法显著更高的翻译质量(例如,使用最小倒数秩)。我们还直接将我们的模型的性能与一种复杂的生成方法进行比较,Haghighi等人使用的匹配典型相关分析(MCCA)算法。(2008)。我们的算法实现了42%的准确性与MCCA的15%。
Bilingual lexicon induction is the task of inducing word translations from monolingual corpora in two languages. In this article we present the most comprehensive analysis of bilingual lexicon induction to date. We present experiments on a wide range of languages and data sizes. We examine translation into English from 25 foreign languages: Albanian, Azeri, Bengali, Bosnian, Bulgarian, Cebuano, Gujarati, Hindi, Hungarian, Indonesian, Latvian, Nepali, Romanian, Serbian, Slovak, Somali, Spanish, Swedish, Tamil, Telugu, Turkish, Ukrainian, Uzbek, Vietnamese, and Welsh. We analyze the behavior of bilingual lexicon induction on low-frequency words, rather than testing solely on high-frequency words, as previous research has done. Low-frequency words are more relevant to statistical machine translation, where systems typically lack translations of rare words that fall outside of their training data. We systematically explore a wide range of features and phenomena that affect the quality of the translations discovered by bilingual lexicon induction. We provide illustrative examples of the highest ranking translations for orthogonal signals of translation equivalence like contextual similarity and temporal similarity. We analyze the effects of frequency and burstiness, and the sizes of the seed bilingual dictionaries and the monolingual training corpora. Additionally, we introduce a novel discriminative approach to bilingual lexicon induction. Our discriminative model is capable of combining a wide variety of features that individually provide only weak indications of translation equivalence. When feature weights are discriminatively set, these signals produce dramatically higher translation quality than previous approaches that combined signals in an unsupervised fashion (e.g., using minimum reciprocal rank). We also directly compare our model's performance against a sophisticated generative approach, the matching canonical correlation analysis (MCCA) algorithm used by Haghighi et al. (2008). Our algorithm achieves an accuracy of 42% versus MCCA's 15%.