Leveraging Parallel Corpora and Existing Wordnets for Automatic Construction of the Slovene Wordnet
Leveraging Parallel Corpora and Existing Wordnets for Automatic Construction of the Slovene Wordnet
复制标题
利用并行语料库和现有词网自动构建斯洛文尼亚词网
DOI:
10.1007/978-3-642-04235-5_31
复制
发表时间:
2009
影响因子:
22.7
通讯作者:
Darja Fišer
中科院分区:
文献类型:
--
作者:
Darja Fišer
The paper reports on a series of experiments conducted in order to test the feasibility of automatically generating synsets for Slovene wordnet. The resources used were the multilingual parallel corpus of George Orwell's Nineteen Eighty-Four and wordnets for several languages. First, the corpus was word-aligned to obtain multilingual lexicons and then these lexicons were compared to the wordnets in various languages in order to disambiguate the entries and attach appropriate synset ids to Slovene entries in the lexicon. Slovene lexicon entries sharing the same attached synset id were then organized into a synset. The results obtained by the different settings in the experiment are evaluated against a manually created gold standard and also checked by hand.