Semantic Similarity Based on Corpus Statistics and Lexical Taxonomy

Semantic Similarity Based on Corpus Statistics and Lexical Taxonomy
复制标题

DOI:
--
复制
发表时间:
1997-08
期刊:
--
影响因子:
--
通讯作者:
Jay J. Jiang;D. Conrath
Jay J. Jiang;D. Conrath
中科院分区:
其他
文献类型:
--
作者:
Jay J. Jiang;D. Conrath

文献摘要

被引文献

相似文献

本文提出了一种新的方法来衡量词和概念之间的语义相似性/距离。它结合了词汇分类结构和语料库统计信息,使得分类法构建的语义空间中节点之间的语义距离可以更好地量化与来自语料库数据的分布分析的计算证据。具体而言,所提出的措施是一种组合的方法,继承了基于边缘的方法的边缘计数方案,然后增强了基于节点的方法的信息内容计算。当测试一个共同的数据集上的词对相似性评级,所提出的方法优于其他计算模型。它给出了最高的相关值(r = 0.828)与基准的基础上人类相似性判断,而上限(r = 0.885)时,观察到人类受试者复制相同的任务。
This paper presents a new approach for measuring semantic similarity/distance between words and concepts. It combines a lexical taxonomy structure with corpus statistical information so that the semantic distance between nodes in the semantic space constructed by the taxonomy can be better quantified with the computational evidence derived from a distributional analysis of corpus data. Specifically, the proposed measure is a combined approach that inherits the edge-based approach of the edge counting scheme, which is then enhanced by the node-based approach of the information content calculation. When tested on a common data set of word pair similarity ratings, the proposed approach outperforms other computational models. It gives the highest correlation value (r = 0.828) with a benchmark based on human similarity judgements, whereas an upper bound (r = 0.885) is observed when human subjects replicate the same task.