Integrating Local and Global Data View for Bilingual Sense Correspondences

Integrating Local and Global Data View for Bilingual Sense Correspondences
复制标题

DOI:
10.1007/978-3-030-15640-4_12
复制
发表时间:
2017-11
期刊:
--
影响因子:
--
通讯作者:
Fumiyo Fukumoto;Yoshimi Suzuki;Attaporn Wangpoonsarp;Meng Ji
Fumiyo Fukumoto;Yoshimi Suzuki;Attaporn Wangpoonsarp;Meng Ji
中科院分区:
其他
文献类型:
--
作者:
Fumiyo Fukumoto;Yoshimi Suzuki;Attaporn Wangpoonsarp;Meng Ji

文献摘要

相似文献

本文提出了一种英汉名词词词典的衔接和双语意义对应的方法。我们使用本地和全局数据视图来识别双语意义对应。局部采用基于简单句子的相似度提取双语名词词。总体而言,对于每个单语词典,我们通过使用具有类别信息的文本语料库来估计领域特定的感官。提取方法是基于词嵌入学习得到的语义相似度。我们合并了这些数据视图。更准确地说,我们为提取的双语单词中的每个名词单词分配了一个意义,保持了领域(类别)的一致性。我们使用WordNet 3.0和EDR日语词典,使用路透社和每日新闻的日语报纸语料库来评估我们的方法。结果表明,本地和全局数据视图的整合提高了整体性能,我们在前1000个双语名词感官中获得了318个。此外,我们发现提取的双语名词意义可以作为机器翻译的词汇资源,使用我们的方法获得的翻译结果优于双语词典,略好于SYSTRANet的翻译结果。
This paper presents a method of linking and creating bilingual sense correspondences between English and Japanese noun word dictionaries. We used local and global data views to identify bilingual sense correspondences. Locally, we extracted bilingual noun words by using simple sentence-based similarity. Globally, for each monolingual dictionary, we estimated domain-specific senses by using a textual corpus having category information. The extraction method is based on the sense similarities which are obtained by word embedding learning. We incorporated these data views. More precisely, we assigned a sense to each noun word of the extracted bilingual words keeping domain (category) consistency. We used the WordNet 3.0 and EDR Japanese dictionaries using Reuters and Mainichi Japanese newspaper corpora to evaluate our method. The results showed that the integration of local and global data views improved overall performance and we obtained 318 within the topmost 1,000 bilingual noun senses. Moreover, we found that the extracted bilingual noun senses can be used as a lexical resource for the machine translation as the translation results obtained by using our method was better than those obtained by a bilingual dictionary and slightly better than the results obtained by SYSTRANet.