Integrating Conflicting Data: The Role of Source Dependence

Integrating Conflicting Data: The Role of Source Dependence
复制标题

DOI:
10.14778/1687627.1687690
复制
发表时间:
2009-08
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
X. Dong;Laure Berti-Équille;D. Srivastava
X. Dong;Laure Berti-Équille;D. Srivastava
中科院分区:
其他
文献类型:
--
作者:
X. Dong;Laure Berti-Équille;D. Srivastava

文献摘要

被引文献

相似文献

许多数据管理应用程序,如设置Web门户、管理企业数据、管理社区数据和共享科学数据,都需要集成来自多个数据源的数据。这些来源中的每一个都提供了一组值,不同的来源通常会提供冲突的值。为了向用户提供高质量的数据,关键是数据集成系统能够解决冲突并发现真正的价值。通常,我们期望一个真值由更多的来源提供,而不是任何一个特定的假值,所以我们可以把大多数来源提供的值作为真值。不幸的是,错误的价值观可以通过复制传播,这使得发现真相变得非常棘手。在本文中,我们考虑了在信息源数量众多,其中一些信息源可能抄袭其他信息源的情况下,如何从相互冲突的信息中找到真实值。我们提出了一种新的方法,考虑了真值发现中数据源之间的依赖关系。直观地说,如果两个数据源提供了大量的公共值,而其中许多值很少由其他数据源提供(例如,特定的假值),那么很可能是一个数据源复制了另一个数据源。我们应用贝叶斯分析来确定信息源之间的依赖关系,并设计了一种迭代检测依赖关系和从冲突信息中发现真相的算法。我们还通过考虑数据源的准确性和值之间的相似性来扩展我们的模型。我们在合成数据和真实数据上的实验表明,我们的算法可以显著提高真相发现的准确性,并且在数据源大量的情况下具有可扩展性。
Many data management applications, such as setting up Web portals, managing enterprise data, managing community data, and sharing scientific data, require integrating data from multiple sources. Each of these sources provides a set of values and different sources can often provide conflicting values. To present quality data to users, it is critical that data integration systems can resolve conflicts and discover true values. Typically, we expect a true value to be provided by more sources than any particular false one, so we can take the value provided by the majority of the sources as the truth. Unfortunately, a false value can be spread through copying and that makes truth discovery extremely tricky. In this paper, we consider how to find true values from conflicting information when there are a large number of sources, among which some may copy from others. We present a novel approach that considers dependence between data sources in truth discovery. Intuitively, if two data sources provide a large number of common values and many of these values are rarely provided by other sources (e.g., particular false values), it is very likely that one copies from the other. We apply Bayesian analysis to decide dependence between sources and design an algorithm that iteratively detects dependence and discovers truth from conflicting information. We also extend our model by considering accuracy of data sources and similarity between values. Our experiments on synthetic data as well as real-world data show that our algorithm can significantly improve accuracy of truth discovery and is scalable when there are a large number of data sources.