Evaluating Word Sense Induction and Disambiguation Methods

Evaluating Word Sense Induction and Disambiguation Methods
复制标题

DOI:
10.1007/s10579-012-9205-0
复制
发表时间:
2013-09
影响因子:
2.7
通讯作者:
Ioannis P. Klapaftis;S. Manandhar
Ioannis P. Klapaftis;S. Manandhar
中科院分区:
计算机科学4区
文献类型:
--
作者:
Ioannis P. Klapaftis;S. Manandhar

文献摘要

被引文献

相似文献

词义归纳 (WSI) 是以无监督的方式识别给定文本中目标词的不同用途(含义)的任务,即不依赖任何外部资源,例如词典或语义标记数据。本文全面描述了 SemEval-2010 WSI 任务以及意义归纳方法的新评估设置。我们的贡献有两个:首先,我们对 Semeval-2010 WSI 任务评估结果进行了详细分析,并找出了当前评估措施的缺点。其次,我们提出了一种新的评估设置,根据目标词的语义分布的偏度来评估参与系统的性能,表明有些方法能够在高度偏斜的分布中远远高于最常见的语义(MFS)基线。
Word Sense Induction (WSI) is the task of identifying the different uses (senses) of a target word in a given text in an unsupervised manner, i.e. without relying on any external resources such as dictionaries or sense-tagged data. This paper presents a thorough description of the SemEval-2010 WSI task and a new evaluation setting for sense induction methods. Our contributions are two-fold: firstly, we provide a detailed analysis of the Semeval-2010 WSI task evaluation results and identify the shortcomings of current evaluation measures. Secondly, we present a new evaluation setting by assessing participating systems’ performance according to the skewness of target words’ distribution of senses showing that there are methods able to perform well above the Most Frequent Sense (MFS) baseline in highly skewed distributions.