Alaska: A Flexible Benchmark for Data Integration Tasks

Alaska: A Flexible Benchmark for Data Integration Tasks
复制标题

阿拉斯加:数据集成任务的灵活基准

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
D. Srivastava
D. Srivastava
中科院分区:
--
文献类型:
--
作者:
Valter Crescenzi;A. D. Angelis;D. Firmani;Maurizio Mazzei;P. Merialdo;Federico Piai;D. Srivastava

文献摘要

被引文献

相似文献

数据集成是数据管理领域长期关注的问题,并且有许多不同的应用,包括商业、科学和政府领域。由于基准测试的可用性不断提高,我们最近在特定的数据集成任务(例如实体解析)中看到了令人瞩目的成果。此类基准测试的一个局限是,它们通常有自己的任务定义,并且可能难以将它们用于复杂的集成流程。因此,对整个数据集成过程的端到端流程进行评估仍然是一个难以实现的目标。在这项工作中,我们提出了阿拉斯加(Alaska),这是第一个基于真实世界数据集的基准测试,可无缝支持数据集成流程中的多个任务(及其变体)。该数据集包含来自71个电子商务网站的约7万个异构产品规格,以及数千个不同的产品属性。我们的基准测试带有分析元数据、一组具有不同特征的预定义用例以及大量人工精心整理的基准事实。我们通过关注两个关键数据集成任务(模式匹配和实体解析)的几种变体来展示我们的基准测试的灵活性。我们的实验表明,我们的基准测试能够对以前难以比较的各种方法进行评估,并能够促进更全面的数据集成解决方案的设计。
Data integration is a long-standing interest of the data management community and has many disparate applications, including business, science and government. We have recently witnessed impressive results in specific data integration tasks, such as Entity Resolution, thanks to the increasing availability of benchmarks. A limitation of such benchmarks is that they typically come with their own task definition and it can be difficult to leverage them for complex integration pipelines. As a result, evaluating end-to-end pipelines for the entire data integration process is still an elusive goal. In this work, we present Alaska, the first benchmark based on real-world dataset to support seamlessly multiple tasks (and their variants) of the data integration pipeline. The dataset consists of ~70k heterogeneous product specifications from 71 e-commerce websites with thousands of different product attributes. Our benchmark comes with profiling meta-data, a set of pre-defined use cases with diverse characteristics, and an extensive manually curated ground truth. We demonstrate the flexibility of our benchmark by focusing on several variants of two crucial data integration tasks, Schema Matching and Entity Resolution. Our experiments show that our benchmark enables the evaluation of a variety of methods that previously were difficult to compare, and can foster the design of more holistic data integration solutions.