A Truth Discovery Approach with Theoretical Guarantee

A Truth Discovery Approach with Theoretical Guarantee
复制标题

DOI:
10.1145/2939672.2939816
复制
发表时间:
2016-08
期刊:
Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Houping Xiao;Jing Gao;Zhaoran Wang;Shiyu Wang;Lu Su;Han Liu
Houping Xiao;Jing Gao;Zhaoran Wang;Shiyu Wang;Lu Su;Han Liu
中科院分区:
其他
文献类型:
--
作者:
Houping Xiao;Jing Gao;Zhaoran Wang;Shiyu Wang;Lu Su;Han Liu

文献摘要

被引文献

相似文献

在信息时代,人们可以很容易地从多个来源收集关于同一组实体的信息,其中冲突是不可避免的。这就引出了一项重要的任务,即发现真相,通过迭代更新真相和来源可靠性来识别真实事实(真相)。然而,在现有的工作中从未讨论过真理的收敛,因此这些真理发现方法的结果没有理论保证。相反,在本文中,我们提出了一个真理发现的方法与理论保证。我们提出了一个随机高斯混合模型(RGMM)来表示多源数据,其中真值是模型参数。我们将源偏置捕获其可靠性程度RGMM制定。然后,将真相发现任务建模为寻求真相的最大似然估计(MLE)。基于期望最大化(EM)技术,我们提出了基于人口的(即,在无限数据的限制下)和基于样本的(即,在有限样本集上)的MLE的解。在理论上,我们证明了在一定条件下,这两个解在极大似然估计周围都是压缩的。实验上,我们评估我们的方法在模拟和真实世界的数据集。实验结果表明,该方法在保证收敛性的前提下,具有较高的识别精度。
In the information age, people can easily collect information about the same set of entities from multiple sources, among which conflicts are inevitable. This leads to an important task, truth discovery, i.e., to identify true facts (truths) via iteratively updating truths and source reliability. However, the convergence to the truths is never discussed in existing work, and thus there is no theoretical guarantee in the results of these truth discovery approaches. In contrast, in this paper we propose a truth discovery approach with theoretical guarantee. We propose a randomized gaussian mixture model (RGMM) to represent multi-source data, where truths are model parameters. We incorporate source bias which captures its reliability degree into RGMM formulation. The truth discovery task is then modeled as seeking the maximum likelihood estimate (MLE) of the truths. Based on expectation-maximization (EM) techniques, we propose population-based (i.e., on the limit of infinite data) and sample-based (i.e., on a finite set of samples) solutions for the MLE. Theoretically, we prove that both solutions are contractive to an ε-ball around the MLE, under certain conditions. Experimentally, we evaluate our method on both simulated and real-world datasets. Experimental results show that our method achieves high accuracy in identifying truths with convergence guarantee.