Unsupervised WSD by Finding the Predominant Sense Using Context as a Dynamic Thesaurus

Unsupervised WSD by Finding the Predominant Sense Using Context as a Dynamic Thesaurus
复制标题

DOI:
10.1007/s11390-010-9385-2
复制
发表时间:
2010-09
影响因子:
0.7
通讯作者:
Javier Tejada-Cárcamo;Hiram Calvo;Alexander Gelbukh;Kazuo Hara
Javier Tejada-Cárcamo;Hiram Calvo;Alexander Gelbukh;Kazuo Hara
中科院分区:
--
文献类型:
--
作者:
Javier Tejada-Cárcamo;Hiram Calvo;Alexander Gelbukh;Kazuo Hara

文献摘要

相似文献

提出并分析了一种无监督的词义消歧方法。我们的工作是基于McCarthyet等人在2004年提出的在整个语料库中寻找每个词的主导意义的方法。他们的最大化算法允许来自分布式词库的加权项(相似词)为每个模糊词义累积得分,即,基于来自与歧义词相关的术语的加权列表的投票来选择具有最高得分的意义。此表是使用林德康提出的分布相似性方法获得的,以获得一个叙词表。在McCarthyet等人的方法中,歧义词的每次出现都使用相同的词库,而不管歧义词出现的上下文。我们的方法占一个词的上下文时,确定一个模糊的词的意义,通过建立一个列表的分布式相似的单词的基础上的语法上下文的模糊的词。我们获得的最高精确度为77.54%的准确度,而在SemCor上测试的原始方法为67.10%。我们还分析了加权词的数量在寻找最常见意义(MFS)和词义消歧任务中的影响,并在几个语料库中进行了实验,以建立词空间模型。
We present and analyze an unsupervised method for Word Sense Disambiguation (WSD). Our work is based on the method presented by McCarthyet al. in 2004 for finding the predominant sense of each word in the entire corpus. Their maximization algorithm allows weighted terms (similar words) from a distributional thesaurus to accumulate a score for each ambiguous word sense, i.e., the sense with the highest score is chosen based on votes from a weighted list of terms related to the ambiguous word. This list is obtained using the distributional similarity method proposed by Lin Dekang to obtain a thesaurus. In the method of McCarthyet al., every occurrence of the ambiguous word uses the same thesaurus, regardless of the context where the ambiguous word occurs. Our method accounts for the context of a word when determining the sense of an ambiguous word by building the list of distributed similar words based on the syntactic context of the ambiguous word. We obtain a top precision of 77.54% of accuracy versus 67.10% of the original method tested on SemCor. We also analyze the effect of the number of weighted terms in the tasks of finding the Most Frecuent Sense (MFS) and WSD, and experiment with several corpora for building the Word Space Model.