A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration

A Bayesian Approach to Discovering Truth from Conflicting Sources for Data Integration
复制标题

DOI:
10.14778/2168651.2168656
复制
发表时间:
2012-02
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Bo Zhao;Benjamin I. P. Rubinstein;J. Gemmell;Jiawei Han
Bo Zhao;Benjamin I. P. Rubinstein;J. Gemmell;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Bo Zhao;Benjamin I. P. Rubinstein;J. Gemmell;Jiawei Han

文献摘要

被引文献

相似文献

在实际的数据集成系统中,被集成的数据源提供有关同一实体的冲突信息是很常见的。因此,数据集成的一个主要挑战是从不同的、有时甚至是相互冲突的来源中获取最完整、最准确的集成记录。我们将这一挑战称为真相发现问题。我们观察到,某些来源通常比其他来源更可靠,因此良好的来源质量模型是解决真相发现问题的关键。在这项工作中,我们提出了一种概率图形模型,可以在没有任何监督的情况下自动推断真实记录和源质量。与以前的方法相比,我们的原则性方法通过对源质量的两个不同方面进行建模,利用两种类型的错误(误报和误报)的生成过程。这样,我们的方法也是第一个旨在合并多值属性类型的方法。我们的方法是可扩展的,因为基于采样的高效推理算法在实践中需要很少的迭代,并且具有线性时间复杂度,并且具有更快的增量变体。对两个现实世界数据集的实验表明,我们的新方法优于现有的解决真相问题的最先进方法。
In practical data integration systems, it is common for the data sources being integrated to provide conflicting information about the same entity. Consequently, a major challenge for data integration is to derive the most complete and accurate integrated records from diverse and sometimes conflicting sources. We term this challenge the truth finding problem. We observe that some sources are generally more reliable than others, and therefore a good model of source quality is the key to solving the truth finding problem. In this work, we propose a probabilistic graphical model that can automatically infer true records and source quality without any supervision. In contrast to previous methods, our principled approach leverages a generative process of two types of errors (false positive and false negative) by modeling two different aspects of source quality. In so doing, ours is also the first approach designed to merge multi-valued attribute types. Our method is scalable, due to an efficient sampling-based inference algorithm that needs very few iterations in practice and enjoys linear time complexity, with an even faster incremental variant. Experiments on two real world datasets show that our new method outperforms existing state-of-the-art approaches to the truth finding problem.