Large scale instance matching via multiple indexes and candidate selection

Large scale instance matching via multiple indexes and candidate selection
复制标题

通过多个索引和候选选择进行大规模实例匹配

DOI:
10.1016/j.knosys.2013.06.004
复制
发表时间:
2013-09-01
影响因子:
8.8
通讯作者:
Tang, Jie
Tang, Jie
中科院分区:
计算机科学1区
文献类型:
--
作者:
Li, Juanzi;Wang, Zhichun;Tang, Jie

文献摘要

被引文献

相似文献

实例匹配旨在发现异构数据源中真实的对象的不同描述之间的联系。随着语义Web的快速发展,尤其是关联数据的快速发展,实例自动匹配已成为实现本体数据共享和集成的基本问题。本体中的对象往往是大规模的,包含数百万甚至数亿个对象。直接应用以往的模式级本体匹配方法是不可行的。本文系统地研究了实例匹配的特点,提出了一种可扩展的、高效的实例匹配方法VMI。VMI为本体实例中不同类型的内容生成多个向量,并使用一组基于倒排索引的规则来获得主要的匹配候选。然后使用用户自定义的属性值来进一步消除不正确的匹配。最后计算匹配候选的相似度作为综合向量距离,提取匹配结果。在OAEI 2009和OAEI 2010的实例跟踪上的实验表明,该方法具有更好的有效性和效率(加速比超过100倍,并且在大多数数据集上的性能(F1得分为+3.0%至5.0%)优于性能最好的Ripperson)。在Linked MDB和DBpedia上的实验表明,VMI可以获得与SILK系统相当的结果(约26,000个质量良好的结果)。(C)2013 Elsevier B.V.保留所有权利。
Instance matching aims to discover the linkage between different descriptions of real objects across heterogeneous data sources. With the rapid development of Semantic Web, especially of the linked data, automatically instance matching has been become the fundamental issue for ontological data sharing and integration. Instances in the ontologies are often in large scale, which contains millions of, or even hundreds of millions objects. Directly applying previous schema level ontology matching methods is infeasible. In this paper, we systematically investigate the characteristics of instance matching, and then propose a scalable and efficient instance matching approach named VMI. VMI generates multiple vectors for different kinds of intained in the ontology instances, and uses a set of inverted indexes based rules to get the primary matching candidates. Then it employs user customized property values to further eliminate the incorrect matchings. Finally the similarities of matching candidates are computed as the integrated vector distances and the matching results are extracted. Experiments on instance track from OAEI 2009 and OAEI 2010 show that the proposed method achieves better effectiveness and efficiency (a speedup of more than 100 times and a bit better performance (+3.0% to 5.0% in terms of F1-score) than top performer RiMOM on most of the datasets). Experiments on Linked MDB and DBpedia show that VMI can obtain comparable results with the SILK system (about 26,000 results with good quality). (C) 2013 Elsevier B.V. All rights reserved.