A Web Search Engine-Based Approach to Measure Semantic Similarity between Words

A Web Search Engine-Based Approach to Measure Semantic Similarity between Words
复制标题

DOI:
10.1109/tkde.2010.172
复制
发表时间:
2011-07
影响因子:
8.9
通讯作者:
Danushka Bollegala;Y. Matsuo;M. Ishizuka
Danushka Bollegala;Y. Matsuo;M. Ishizuka
中科院分区:
计算机科学2区
文献类型:
--
作者:
Danushka Bollegala;Y. Matsuo;M. Ishizuka

文献摘要

被引文献

相似文献

度量词之间的语义相似性是Web上各种任务的重要组成部分,例如关系提取,社区挖掘,文档聚类和自动元数据提取。尽管语义相似性度量在这些应用中是有用的,但准确地测量两个词(或实体)之间的语义相似性仍然是一项具有挑战性的任务。我们提出了一个经验的方法来估计语义相似性,使用网页计数和文本片段从Web搜索引擎检索两个字。具体来说,我们定义了各种词同现措施使用页数和整合那些从文本片段中提取的词汇模式。为了识别两个给定单词之间存在的大量语义关系,我们提出了一种新的模式提取算法和模式聚类算法。使用支持向量机学习基于页面计数的共现度量和词汇模式聚类的最佳组合。所提出的方法优于各种基线和先前提出的基于Web的语义相似性措施的三个基准数据集显示出与人类评级的高度相关性。此外,该方法显着提高了社区挖掘任务的准确性。
Measuring the semantic similarity between words is an important component in various tasks on the web such as relation extraction, community mining, document clustering, and automatic metadata extraction. Despite the usefulness of semantic similarity measures in these applications, accurately measuring semantic similarity between two words (or entities) remains a challenging task. We propose an empirical method to estimate semantic similarity using page counts and text snippets retrieved from a web search engine for two words. Specifically, we define various word co-occurrence measures using page counts and integrate those with lexical patterns extracted from text snippets. To identify the numerous semantic relations that exist between two given words, we propose a novel pattern extraction algorithm and a pattern clustering algorithm. The optimal combination of page counts-based co-occurrence measures and lexical pattern clusters is learned using support vector machines. The proposed method outperforms various baselines and previously proposed web-based semantic similarity measures on three benchmark data sets showing a high correlation with human ratings. Moreover, the proposed method significantly improves the accuracy in a community mining task.