VDoc+: a virtual document based approach for matching large ontologies using MapReduce

VDoc+: a virtual document based approach for matching large ontologies using MapReduce
复制标题

DOI:
10.1631/jzus.c1101007
复制
发表时间:
2012-04
期刊:
Journal of Zhejiang University SCIENCE C
影响因子:
--
通讯作者:
Hang Zhang;Wei Hu;Yuzhong Qu
Hang Zhang;Wei Hu;Yuzhong Qu
中科院分区:
其他
文献类型:
--
作者:
Hang Zhang;Wei Hu;Yuzhong Qu

文献摘要

相似文献

许多本体已经发布在语义Web上,以供共享来描述资源。其中,现实领域的大型本体在提出本体匹配(OM)等语义技术时存在可扩展性问题。这要么是运行时间太长,要么是对运行环境有很强的假设。针对这一问题,本文在MapReduce框架和虚拟文档技术的基础上,提出了一种基于MapReduce的三阶段大型本体匹配方法V-Doc+。具体地说,在第一阶段执行两个MapReduce过程,以分别提取命名实体(类、属性和实例)和空白节点的文本描述。在第二阶段,将提取的描述与资源描述框架(RDF)图中的邻居交换,以构建虚拟文档。此提取过程还受益于基于MapReduce的实现。第三阶段提出了一种基于词权重的分词方法,利用词频-逆文频(TF-IDF)模型进行并行相似度计算。在两个大规模真实数据集和OAEI的基准测试床上的实验结果表明,该方法在准确率和召回率损失较小的情况下,显著减少了运行时间。
Many ontologies have been published on the Semantic Web, to be shared to describe resources. Among them, large ontologies of real-world areas have the scalability problem in presenting semantic technologies such as ontology matching (OM). This either suffers from too long run time or has strong hypotheses on the running environment. To deal with this issue, we propose a three-stage MapReduce-based approach V-Doc+ for matching large ontologies, based on the MapReduce framework and virtual document technique. Specifically, two MapReduce processes are performed in the first stage to extract the textual descriptions of named entities (classes, properties, and instances) and blank nodes, respectively. In the second stage, the extracted descriptions are exchanged with neighbors in Resource Description Framework (RDF) graphs to construct virtual documents. This extraction process also benefits from the MapReduce-based implementation. A word-weight-based partitioning method is proposed in the third stage to conduct parallel similarity calculation using the term frequency-inverse document frequency (TF-IDF) model. Experimental results on two large-scale real datasets and the benchmark testbed from Ontology Alignment Evaluation Initiative (OAEI) are reported, showing that the proposed approach significantly reduces the run time with minor loss in precision and recall.