Large scale instance matching via multiple indexes and candidate selection
Large scale instance matching via multiple indexes and candidate selection
复制标题
通过多个索引和候选选择进行大规模实例匹配
DOI:
10.1016/j.knosys.2013.06.004
复制
发表时间:
2013-09-01
影响因子:
8.8
通讯作者:
Tang, Jie
中科院分区:
文献类型:
--
作者:
Li, Juanzi;Wang, Zhichun;Tang, Jie
Instance matching aims to discover the linkage between different descriptions of real objects across heterogeneous data sources. With the rapid development of Semantic Web, especially of the linked data, automatically instance matching has been become the fundamental issue for ontological data sharing and integration. Instances in the ontologies are often in large scale, which contains millions of, or even hundreds of millions objects. Directly applying previous schema level ontology matching methods is infeasible. In this paper, we systematically investigate the characteristics of instance matching, and then propose a scalable and efficient instance matching approach named VMI. VMI generates multiple vectors for different kinds of intained in the ontology instances, and uses a set of inverted indexes based rules to get the primary matching candidates. Then it employs user customized property values to further eliminate the incorrect matchings. Finally the similarities of matching candidates are computed as the integrated vector distances and the matching results are extracted. Experiments on instance track from OAEI 2009 and OAEI 2010 show that the proposed method achieves better effectiveness and efficiency (a speedup of more than 100 times and a bit better performance (+3.0% to 5.0% in terms of F1-score) than top performer RiMOM on most of the datasets). Experiments on Linked MDB and DBpedia show that VMI can obtain comparable results with the SILK system (about 26,000 results with good quality). (C) 2013 Elsevier B.V. All rights reserved.