A cross-lingual similarity measure for detecting biomedical term translations.

A cross-lingual similarity measure for detecting biomedical term translations.
复制标题

DOI:
10.1371/journal.pone.0126196
复制
发表时间:
2015
期刊:
影响因子:
3.7
通讯作者:
Ananiadou S
Ananiadou S
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Bollegala D;Kontonatsios G;Ananiadou S

文献摘要

参考文献

被引文献

相似文献

生物医学术语等技术术语的双语词典是机器翻译系统的重要资源,也是希望理解外语概念的人的重要资源。通常,生物医学术语最初是用英语提出的,后来被手动翻译成其他语言。尽管有大量的单语生物医学术语词典,但这些术语词典中只有一小部分被翻译成其他语言。人工编制大型双语词典的技术领域是一项具有挑战性的任务,因为它是很难找到足够多的双语专家。我们提出了一个跨语言的相似性度量检测最相似的翻译候选人的一种语言(源)从另一种语言(目标)指定的生物医学术语。具体地,使用两种类型的特征来表示语言中的生物医学术语:(a)由从所考虑的术语中提取的字符n-gram组成的内在特征,以及(B)由从所考虑的术语周围的上下文窗口中提取的一元语法和二元语法组成的外在特征。我们提出了一个跨语言的相似性度量使用这些特征类型。首先,为了降低每种语言的特征空间的维数,我们提出了原型向量投影(PVP)-一种非负的低维向量投影方法。其次,我们提出了一种方法来学习的源和目标语言的特征空间之间的映射,使用偏最小二乘回归(PLSR)。该方法只需要少量的训练实例来学习跨语言的相似性度量。所提出的PVP方法优于流行的降维方法,如奇异值分解(SVD)和非负矩阵分解(NMF)在最近邻预测任务。此外,我们的实验结果,包括几个语言对,如英语-法语,英语-西班牙语,英语-希腊语,英语-日语表明,该方法优于其他几个特征投影方法在生物医学术语翻译预测任务。
Bilingual dictionaries for technical terms such as biomedical terms are an important resource for machine translation systems as well as for humans who would like to understand a concept described in a foreign language. Often a biomedical term is first proposed in English and later it is manually translated to other languages. Despite the fact that there are large monolingual lexicons of biomedical terms, only a fraction of those term lexicons are translated to other languages. Manually compiling large-scale bilingual dictionaries for technical domains is a challenging task because it is difficult to find a sufficiently large number of bilingual experts. We propose a cross-lingual similarity measure for detecting most similar translation candidates for a biomedical term specified in one language (source) from another language (target). Specifically, a biomedical term in a language is represented using two types of features: (a) intrinsic features that consist of character n-grams extracted from the term under consideration, and (b) extrinsic features that consist of unigrams and bigrams extracted from the contextual windows surrounding the term under consideration. We propose a cross-lingual similarity measure using each of those feature types. First, to reduce the dimensionality of the feature space in each language, we propose prototype vector projection (PVP)—a non-negative lower-dimensional vector projection method. Second, we propose a method to learn a mapping between the feature spaces in the source and target language using partial least squares regression (PLSR). The proposed method requires only a small number of training instances to learn a cross-lingual similarity measure. The proposed PVP method outperforms popular dimensionality reduction methods such as the singular value decomposition (SVD) and non-negative matrix factorization (NMF) in a nearest neighbor prediction task. Moreover, our experimental results covering several language pairs such as English–French, English–Spanish, English–Greek, and English–Japanese show that the proposed method outperforms several other feature projection methods in biomedical term translation prediction tasks.
DOI: 10.1007/11562214_55
发表时间: 2005-01-01
期刊: NATURAL LANGUAGE PROCESSING - IJCNLP 2005, PROCEEDINGS
影响因子: --
作者:
Bollegala, D;Okazaki, N;Ishizuka, M
通讯作者: Ishizuka, M
DOI: 10.1016/j.csda.2008.01.011
发表时间: 2008-04-15
影响因子: 1.8
作者:
Ding, Chris;Li, Tao;Peng, Wei
通讯作者: Peng, Wei