Parallel and Streaming Truth Discovery in Large-Scale Quantitative Crowdsourcing

Parallel and Streaming Truth Discovery in Large-Scale Quantitative Crowdsourcing
复制标题

DOI:
10.1109/tpds.2016.2515092
复制
发表时间:
2016-10
影响因子:
5.3
通讯作者:
W. Ouyang;Lance M. Kaplan;Alice Toniolo;M. Srivastava;T. Norman
W. Ouyang;Lance M. Kaplan;Alice Toniolo;M. Srivastava;T. Norman
中科院分区:
计算机科学2区
文献类型:
--
作者:
W. Ouyang;Lance M. Kaplan;Alice Toniolo;M. Srivastava;T. Norman

文献摘要

被引文献

相似文献

为了实现可靠的众包应用,开发能够从各种信息源提供的可能嘈杂和冲突的主张中自动发现真相的算法是非常重要的。为了处理涉及大数据或流数据的众包应用,理想的真相发现算法不仅应该是有效的,而且是可扩展的。然而,相对于定量众包应用,如对象计数和百分比注释,现有的真相发现算法是不同时有效和可扩展的。它们要么在分类众包中解决真相发现问题,要么执行无法扩展的批处理。在本文中,我们提出了新的并行和流式真相发现算法的定量众包应用。通过在真实数据集和合成数据集上的大量实验,我们证明了:1)这两种算法都是非常有效的; 2)并行算法可以有效地在大数据集上进行真值发现; 3)流式算法以增量方式处理数据,可以有效地在大数据集和数据流中进行真值发现。
To enable reliable crowdsourcing applications, it is of great importance to develop algorithms that can automatically discover the truths from possibly noisy and conflicting claims provided by various information sources. In order to handle crowdsourcing applications involving big or streaming data, a desirable truth discovery algorithm should not only be effective, but also be scalable. However, with respect to quantitative crowdsourcing applications such as object counting and percentage annotation, existing truth discovery algorithms are not simultaneously effective and scalable. They either address truth discovery in categorical crowdsourcing or perform batch processing that does not scale. In this paper, we propose new parallel and streaming truth discovery algorithms for quantitative crowdsourcing applications. Through extensive experiments on real-world and synthetic datasets, we demonstrate that 1) both of them are quite effective, 2) the parallel algorithm can efficiently perform truth discovery on large datasets, and 3) the streaming algorithm processes data incrementally, and it can efficiently perform truth discovery both on large datasets and in data streams.