Computing semantic similarity based on novel models of semantic representation using Wikipedia

Computing semantic similarity based on novel models of semantic representation using Wikipedia
复制标题

使用维基百科基于新颖的语义表示模型计算语义相似度

DOI:
10.1016/j.ipm.2018.07.002
复制
发表时间:
2018-11-01
影响因子:
8.6
通讯作者:
Jiang, Yuncheng
Jiang, Yuncheng
中科院分区:
计算机科学1区
文献类型:
--
作者:
Qu, Rong;Fang, Yongyi;Jiang, Yuncheng

文献摘要

被引文献

相似文献

计算概念之间的语义相似度(SS)是自然语言处理和人工智能等许多领域最关键的问题之一。多年来,通过利用不同的知识资源,已经提出了几种SS测量方法。维基百科提供了一个大型的独立于领域的百科全书存储库和一个用于计算概念之间的 SS 的语义网络。传统的基于特征的度量依赖于不同属性的线性组合,有两个主要局限性:信息不足和语义信息丢失。在本文中,我们利用信息内容(IC)和概念特征提出了几种混合SS测量方法,避免了上述限制。考虑将离散属性集成到一个组件中,我们提出了两种语义表示模型,称为 CORM 和 CARM。然后,我们基于这些模型计算SS,并将类别的IC作为SS测量的补充。该评估基于几个广泛使用的基准和我们自己开发的基准,维持了人类判断的直觉。总之,我们的方法在确定概念之间的 SS 方面更有效,并且比以前的方法(例如 Word2Vec 和 NASARI)具有更好的人类相关性。
Computing Semantic Similarity (SS) between concepts is one of the most critical issues in many domains such as Natural Language Processing and Artificial Intelligence. Over the years, several SS measurement methods have been proposed by exploiting different knowledge resources. Wikipedia provides a large domain-independent encyclopedic repository and a semantic network for computing SS between concepts. Traditional feature-based measures rely on linear combinations of different properties with two main limitations, the insufficient information and the loss of semantic information. In this paper, we propose several hybrid SS measurement approaches by using the Information Content (IC) and features of concepts, which avoid the limitations introduced above. Considering integrating discrete properties into one component, we present two models of semantic representation, called CORM and CARM. Then, we compute SS based on these models and take the IC of categories as a supplement of SS measurement. The evaluation, based on several widely used benchmarks and a benchmark developed by ourselves, sustains the intuitions with respect to human judgments. In summary, our approaches are more efficient in determining SS between concepts and have a better human correlation than previous methods such as Word2Vec and NASARI.