A hybrid approach for measuring semantic similarity based on IC-weighted path distance in WordNet

A hybrid approach for measuring semantic similarity based on IC-weighted path distance in WordNet
复制标题

DOI:
10.1007/s10844-017-0479-y
复制
发表时间:
2018-08
影响因子:
3.4
通讯作者:
Yuanyuan Cai;Qingchuan Zhang;W. Lu;Xiaoping Che
Yuanyuan Cai;Qingchuan Zhang;W. Lu;Xiaoping Che
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yuanyuan Cai;Qingchuan Zhang;W. Lu;Xiaoping Che

文献摘要

被引文献

相似文献

作为一种有价值的文本理解工具,语义相似性度量在自然语言处理、信息检索、计算语言学和人工智能等领域提供了基于区分语义的应用。现有的研究大多采用WordNet等结构化分类法来探索词汇语义关系,然而,计算精度的提高仍然是它们面临的挑战。针对这一问题,本文提出了一种基于混合WordNet的概念语义相似度度量方法CSSM-ICSP,该方法利用概念的信息量来加权概念之间的最短路径距离。为了提高IC计算的性能,我们还提出了一种新的概念内在IC模型,其中考虑了WordNet结构中涉及的各种语义属性。此外,我们总结和分类了以前基于WordNet的方法的技术特点,并在不同的基准上对照这些方法对我们的方法进行了评估。实验结果表明,本文提出的IC模型和相似度检测方法在语义相似度度量方面与其他方法具有可比性,甚至更好。
As a valuable tool for text understanding, semantic similarity measurement enables discriminative semantic-based applications in the fields of natural language processing, information retrieval, computational linguistics and artificial intelligence. Most of the existing studies have used structured taxonomies such as WordNet to explore the lexical semantic relationship, however, the improvement of computation accuracy is still a challenge for them. To address this problem, in this paper, we propose a hybrid WordNet-based approach CSSM-ICSP to measuring concept semantic similarity, which leverage the information content(IC) of concepts to weight the shortest path distance between concepts. To improve the performance of IC computation, we also develop a novel model of the intrinsic IC of concepts, where a variety of semantic properties involved in the structure of WordNet are taken into consideration. In addition, we summarize and classify the technical characteristics of previous WordNet-based approaches, as well as evaluate our approach against these approaches on various benchmarks. The experimental results of the proposed approaches are more correlated with human judgment of similarity in term of the correlation coefficient, which indicates that our IC model and similarity detection approach are comparable or even better for semantic similarity measurement as compared to others.