GeoLink Cruises: A Non-Synthetic Benchmark for Co-Reference Resolution on Knowledge Graphs

GeoLink Cruises: A Non-Synthetic Benchmark for Co-Reference Resolution on Knowledge Graphs
复制标题

DOI:
10.1145/3340531.3412770
复制
发表时间:
2020-10
期刊:
Proceedings of the 29th ACM International Conference on Information & Knowledge Management
影响因子:
--
通讯作者:
Reihaneh Amini;Lu Zhou;P. Hitzler
Reihaneh Amini;Lu Zhou;P. Hitzler
中科院分区:
其他
文献类型:
--
作者:
Reihaneh Amini;Lu Zhou;P. Hitzler

文献摘要

相似文献

十多年来,共指消解系统已经被开发出来,以便在不同链接数据集和知识图的实例之间找到简单的1对1等价映射(sameAs关系)。实例匹配系统的比较评估可以告诉我们这些系统在人工基准或真实数据挑战方面的性能。然而,缺乏用于评估这些系统的真实的数据是目前的一个瓶颈。在本文中,我们建议使用的Cruise实体的GeoLink数据存储库作为一个现实世界的实例匹配基准链接的数据和知识图。GeoLink项目汇集了与地球科学研究有关的七个数据集。GeoLink的本体(T-box)和实例数据(A-box)都比当前的基准大得多,并且它们具有特别有趣的挑战,例如地理空间和时间数据。我们在这里提出的基准测试由GeoLink中的两个真实世界数据集组成,称为R2 R数据和BCO-DMO,其中包括手动策划的owl:sameAs链接,这两个数据集的900多个Cruise实体之间。参考对齐由来自不同机构的领域专家讨论并生成,并以对齐API格式表示。
Since over a decade coreference resolution systems have been developed in order to find simple 1-to-1 equivalent mapping (sameAs relations) between instances of different linked datasets and knowledge graphs. Comparative evaluations of instance matching systems can inform us about the performance of such systems regarding artificial benchmarks or real-world data challenges. However, the lack of real data for evaluating these systems is currently a bottleneck. In this paper, we propose the use of the Cruise entities in the GeoLink data repository as a real-world instance matching benchmark for linked data and knowledge graphs. The GeoLink project has brought together seven datasets related to geoscience research. Both the ontology (T-box) and the instance data (A-box) of GeoLink are significantly larger than current benchmarks, and they have particularly interesting challenges, such as geospatial and temporal data. The benchmark we propose here consists of two real-world datasets in GeoLink called R2R data and BCO-DMO which includes manual curated owl:sameAs links between more than 900 Cruise entities of these two datasets. The reference alignment was discussed and generated by domain experts from different institutions and is expressed in the Alignment API format.