A classification approach for detecting cross-lingual biomedical term translations

A classification approach for detecting cross-lingual biomedical term translations
复制标题

DOI:
10.1017/s1351324915000431
复制
发表时间:
2017-01-01
影响因子:
2.5
通讯作者:
Bollegala, D.
Bollegala, D.
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hakami, H.;Bollegala, D.

文献摘要

被引文献

相似文献

技术术语的翻译是机器翻译中的一个重要问题。特别是,在高度专业化的领域,如生物学或医学,很难找到双语专家来注释足够的跨语言文本,以训练机器翻译系统。此外,生物医学界不断产生新术语,这使得翻译词典很难为所有感兴趣的语言对保持最新。给定一种语言(源语言)的生物医学术语,我们提出了一种方法,用于检测其在不同的语言(目标语言)的翻译。具体来说,我们训练一个二元分类器来确定两个用两种语言编写的生物医学术语是否是翻译。由于源语言和目标语言之间缺乏共同特征,训练这样的分类器通常是复杂的。我们提出了几种特征空间拼接方法来成功地克服这个问题。此外,我们研究了上下文和字符的n-gram功能检测术语翻译的有效性。使用标准数据集进行的生物医学术语翻译实验表明,该方法优于几个竞争的基线方法的平均平均精度和top-k翻译精度。
Finding translations for technical terms is an important problem in machine translation. In particular, in highly specialized domains such as biology or medicine, it is difficult to find bilingual experts to annotate sufficient cross-lingual texts in order to train machine translation systems. Moreover, new terms are constantly being generated in the biomedical community, which makes it difficult to keep the translation dictionaries up to date for all language pairs of interest. Given a biomedical term in one language (source language), we propose a method for detecting its translations in a different language (target language). Specifically, we train a binary classifier to determine whether two biomedical terms written in two languages are translations. Training such a classifier is often complicated due to the lack of common features between the source and target languages. We propose several feature space concatenation methods to successfully overcome this problem. Moreover, we study the effectiveness of contextual and character n-gram features for detecting term translations. Experiments conducted using a standard dataset for biomedical term translation show that the proposed method outperforms several competitive baseline methods in terms of mean average precision and top-k translation accuracy.