A Probabilistic Model for Estimating Real-valued Truth from Conflicting Sources

A Probabilistic Model for Estimating Real-valued Truth from Conflicting Sources
复制标题

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Bo Zhao;Jiawei Han
Bo Zhao;Jiawei Han
中科院分区:
其他
文献类型:
--
作者:
Bo Zhao;Jiawei Han

文献摘要

被引文献

相似文献

数据集成中的一项重要任务是从来自多个来源的噪声和冲突的数据记录中识别真相,即真相发现问题。以前,已经提出了几种方法来解决这个问题,方法是同时学习来源的质量和真相。然而,所有这些方法主要是为处理分类数据而非数值数据而设计的。而在实践中,数字数据不仅无处不在,而且价值很高,例如价格、天气、人口普查、民意调查、经济统计等。由于数字数据的特点,数字数据的质量问题甚至可能比分类数据更常见和严重。因此,在这项工作中,我们提出了一种专门为处理数值数据而设计的新的求真方法。基于贝叶斯概率模型,我们的方法可以有原则地利用数字数据的特征,在建模来源质量、真实性和声明值之间的依赖关系时。在两个真实世界的数据集上的实验表明,我们的新方法比现有的最先进的方法性能更好。
One important task in data integration is to identify truth from noisy and conflicting data records collected from multiple sources, i.e., the truth finding problem. Previously, several methods have been proposed to solve this problem by simultaneously learning the quality of sources and the truth. However, all those methods are mainly designed for handling categorical data but not numerical data. While in practice, numerical data is not only ubiquitous but also of high value, e.g. price, weather, census, polls, economic statistics, etc. Quality issues on numerical data can also be even more common and severe than categorical data due to its characteristics. Therefore, in this work we propose a new truth-finding method specially designed for handling numerical data. Based on Bayesian probabilistic models, our method can leverage the characteristics of numerical data in a principled way, when modeling the dependencies among source quality, truth, and claimed values. Experiments on two real world datasets show that our new method outperforms existing state-of-the-art approaches.