Leveraging Parallel Corpora and Existing Wordnets for Automatic Construction of the Slovene Wordnet

Leveraging Parallel Corpora and Existing Wordnets for Automatic Construction of the Slovene Wordnet
复制标题

利用并行语料库和现有词网自动构建斯洛文尼亚词网

DOI:
10.1007/978-3-642-04235-5_31
复制
发表时间:
2009
影响因子:
22.7
通讯作者:
Darja Fišer
Darja Fišer
中科院分区:
计算机科学3区
文献类型:
--
作者:
Darja Fišer

文献摘要

被引文献

相似文献

本文报告了一系列的实验,以测试的可行性,自动生成同义词斯洛文尼亚语的wordnet。使用的资源是乔治奥威尔的《一九八四》的多语言平行语料库和几种语言的词汇网。首先,语料库的词对齐,以获得多语种的词典,然后这些词典进行比较,以消除歧义的条目,并附上适当的同义词集ID斯洛文尼亚的词汇中的条目在各种语言的词网。斯洛文尼亚语词典条目共享相同的附加同义词集ID,然后组织成一个同义词集。实验中通过不同设置获得的结果将根据手动创建的金标准进行评估,并手动检查。
The paper reports on a series of experiments conducted in order to test the feasibility of automatically generating synsets for Slovene wordnet. The resources used were the multilingual parallel corpus of George Orwell's Nineteen Eighty-Four and wordnets for several languages. First, the corpus was word-aligned to obtain multilingual lexicons and then these lexicons were compared to the wordnets in various languages in order to disambiguate the entries and attach appropriate synset ids to Slovene entries in the lexicon. Slovene lexicon entries sharing the same attached synset id were then organized into a synset. The results obtained by the different settings in the experiment are evaluated against a manually created gold standard and also checked by hand.