Improving the extraction of bilingual terminology from Wikipedia

Improving the extraction of bilingual terminology from Wikipedia
复制标题

DOI:
10.1145/1596990.1596995
复制
发表时间:
2009-10
期刊:
ACM Trans. Multim. Comput. Commun. Appl.
影响因子:
--
通讯作者:
M. Erdmann;Kotaro Nakayama;T. Hara;S. Nishio
M. Erdmann;Kotaro Nakayama;T. Hara;S. Nishio
中科院分区:
其他
文献类型:
--
作者:
M. Erdmann;Kotaro Nakayama;T. Hara;S. Nishio

文献摘要

被引文献

相似文献

双语词典的自动构建研究已经取得了令人瞩目的成果。双语词典通常由平行语料库构建,但由于这些语料库仅适用于选定的文本域和语言对,因此也正在探索其他资源的潜力。在这篇文章中,我们想进一步探讨使用维基百科作为双语术语提取语料库的想法。我们提出了一种方法,从不同类型的维基百科链接信息中提取术语翻译对。之后,在人工标记的训练数据的特征上训练的SVM分类器确定看不见的术语翻译对的正确性。
Research on the automatic construction of bilingual dictionaries has achieved impressive results. Bilingual dictionaries are usually constructed from parallel corpora, but since these corpora are available only for selected text domains and language pairs, the potential of other resources is being explored as well. In this article, we want to further pursue the idea of using Wikipedia as a corpus for bilingual terminology extraction. We propose a method that extracts term-translation pairs from different types of Wikipedia link information. After that, an SVM classifier trained on the features of manually labeled training data determines the correctness of unseen term-translation pairs.