A Relational Model of Semantic Similarity between Words using Automatically Extracted Lexical Pattern Clusters from the Web

A Relational Model of Semantic Similarity between Words using Automatically Extracted Lexical Pattern Clusters from the Web
复制标题

DOI:
10.3115/1699571.1699617
复制
发表时间:
2009-08
期刊:
--
影响因子:
--
通讯作者:
Danushka Bollegala;Y. Matsuo;M. Ishizuka
Danushka Bollegala;Y. Matsuo;M. Ishizuka
中科院分区:
其他
文献类型:
--
作者:
Danushka Bollegala;Y. Matsuo;M. Ishizuka

文献摘要

被引文献

相似文献

语义相似性是一个核心概念,涵盖人工智能、自然语言处理、认知科学和心理学等众多领域。准确测量单词之间的语义相似度对于文档聚类、信息检索和同义词提取等各种任务至关重要。我们利用单词之间存在的语义关系提出了一种新的语义相似度模型。给定两个单词,首先,我们使用自动提取的词汇模式簇来表示这些单词之间的语义关系。接下来,使用马哈拉诺比斯距离度量计算两个单词之间的语义相似度。我们将所提出的相似性度量与之前在 Miller-Charles 基准数据集和 WordSimilarity-353 集合上提出的语义相似性度量进行了比较。所提出的方法优于所有现有的基于网络的语义相似性度量,在 Millet-Charles 数据集上实现了 0.867 的皮尔逊相关系数。
Semantic similarity is a central concept that extends across numerous fields such as artificial intelligence, natural language processing, cognitive science and psychology. Accurate measurement of semantic similarity between words is essential for various tasks such as, document clustering, information retrieval, and synonym extraction. We propose a novel model of semantic similarity using the semantic relations that exist among words. Given two words, first, we represent the semantic relations that hold between those words using automatically extracted lexical pattern clusters. Next, the semantic similarity between the two words is computed using a Mahalanobis distance measure. We compare the proposed similarity measure against previously proposed semantic similarity measures on Miller-Charles benchmark dataset and WordSimilarity-353 collection. The proposed method outperforms all existing web-based semantic similarity measures, achieving a Pearson correlation coefficient of 0.867 on the Millet-Charles dataset.