Semantic similarity measures for formal concept analysis using linked data and WordNet

Semantic similarity measures for formal concept analysis using linked data and WordNet
复制标题

使用链接数据和 WordNet 进行形式概念分析的语义相似性度量

DOI:
10.1007/s11042-019-7150-2
复制
发表时间:
2019-07-01
影响因子:
3.6
通讯作者:
Qu, Rong
Qu, Rong
中科院分区:
计算机科学4区
文献类型:
--
作者:
Jiang, Yuncheng;Yang, Mingxuan;Qu, Rong

文献摘要

被引文献

相似文献

形式概念分析(FCA)是一个应用数学领域,其根源在于序论,特别是完备格理论。它不仅是一种数据分析和知识表示的方法,而且是一种概念形成和学习的形式化公式。在过去的20多年里,FCA得到了广泛的研究。本文分析了FCA中相似性度量的研究现状和存在的问题。针对现有方法的不足,提出了一种基于关联数据和WordNet的FCA语义相似度度量方法。我们的目标是开发一种全自动的方法,不需要预定义的领域本体,并且可以独立于领域使用,在FCA中需要语义相似性度量的应用程序中。为了实现FCA的语义相似度评估,首先利用WordNet将关联数据中资源(或实体)的相似度评估方法扩展到语义案例。此外,我们针对FCA概念和概念格分别提出了两种语义相似度度量方法(即上下文无关方法和上下文感知方法)。与FCA中已有的相似性度量方法相比,该方法利用可能性理论的概念来确定相似区间的上下界。最后,我们通过在真实数据集上的应用对所提出的相似性评估方法进行了评估。
Formal Concept Analysis (FCA) is a field of applied mathematics with its roots in order theory, in particular the theory of complete lattices. It is not only a method for data analysis and knowledge representation, but also a formal formulation for concept formation and learning. Over the past 20 years, FCA has been widely studied. In this paper, the current research progresses and the existing problems of similarity measures in FCA are analyzed. To address the drawbacks of the existing methods, we propose a kind of novel semantic similarity measure for FCA by using Linked Data and WordNet. We aim to develop a method that is fully automatic without requiring predefined domain ontologies and can be used independently of the domain in applications requiring semantic similarity measures in FCA. To realize the semantic similarity estimation for FCA, we firstly extend the similarity assessment methods for resources (or entities) in Linked Data into semantic cases by using WordNet. Furthermore, we propose two kinds of semantic similarity measures (i.e., context-free method and context-aware method) for FCA concepts and concept lattices, respectively. Compared with the existing similarity measure methods in FCA, the proposed approach uses concept of possibility theory to determine lower and upper bounds of similarity intervals. Finally, we evaluate the proposed similarity assessment approaches by applying them to real-worlds datasets.