Data veracity estimation with ensembling truth discovery methods

Data veracity estimation with ensembling truth discovery methods
复制标题

使用集成真理发现方法估计数据准确性

DOI:
--
复制
发表时间:
2015
期刊:
2015 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Laure Berti
Laure Berti
中科院分区:
--
文献类型:
--
作者:
Laure Berti

文献摘要

参考文献

被引文献

相似文献

数据准确性的估计被认为是大数据的重大挑战之一。通常,真相发现的目标是确定多源、冲突数据的准确性,并作为输出返回每个数据值的准确性标签和置信度分数,以及声明该数据的每个来源的可信度分数。尽管已经提出了大量方法,但不太可能有一种技术在所有数据集中支配所有其他技术。此外,这些方法的性能评估完全取决于标记的地面真实数据(即,其准确性已被手动检查的数据)的可用性。在大数据背景下,获取完整的地面实况数据是遥不可及的。在本文中,我们提出了一种集成方法,可以缓解方法选择和地面实况数据稀疏性两个问题。我们的方法结合了一组真相发现方法的结果,初步实验表明,当使用地面真相数据样本时,它比单一方法提高了质量性能。
Estimation of data veracity is recognized as one of the grand challenges of big data. Typically, the goal of truth discovery is to determine the veracity of multi-source, conflicting data and return, as outputs, a veracity label and a confidence score for each data value, along with the trustworthiness score of each source claiming it. Although a plethora of methods has been proposed, it is unlikely a technique dominates all others across all data sets. Furthermore, the performance evaluation of the methods entirely depends on the availability of labeled ground truth data (i.e., data whose veracity has been manually checked). In the context of Big Data, acquiring the complete ground truth data is out-of-reach. In this paper, we propose an ensembling method that mitigates the two problems of method selection and ground truth data sparsity. Our approach combines the results of a set of truth discovery methods and preliminary experiments suggest that it improves the quality performance over the single methods when samples of ground truth data are used.
DOI: 10.1007/3-540-45014-9
发表时间: 2000-06
期刊: --
影响因子: --
作者:
Thomas G. Dietterich
通讯作者: Thomas G. Dietterich