Iterative Learning for Reliable Crowdsourcing Systems

Iterative Learning for Reliable Crowdsourcing Systems
复制标题

DOI:
--
复制
发表时间:
2011-12
期刊:
--
影响因子:
--
通讯作者:
David R Karger;Sewoong Oh;Devavrat Shah
David R Karger;Sewoong Oh;Devavrat Shah
中科院分区:
其他
文献类型:
--
作者:
David R Karger;Sewoong Oh;Devavrat Shah

文献摘要

被引文献

相似文献

众包系统,其中任务被电子地分配给许多“信息件工作者”,已经成为一种有效的范例,用于在诸如图像分类、数据输入、光学字符识别、推荐和校对等领域中的大规模问题的人力解决。由于这些低收入的工人可能不可靠,几乎所有的众包都必须设计方案来增加对他们答案的信心,通常是通过多次分配每个任务并以某种方式组合答案,例如多数投票。在本文中,我们考虑了这种众包任务的一般模型,并提出了最小化总价格的问题(即,任务分配的数量),其必须被支付以实现目标总体可靠性。我们给出了一个新的算法来决定哪些任务分配给哪些工人,并从工人的答案推断正确的答案。我们表明,我们的算法显着优于多数表决,事实上,是渐近最优的,通过比较的甲骨文,知道每个工人的可靠性。
Crowdsourcing systems, in which tasks are electronically distributed to numerous "information piece-workers", have emerged as an effective paradigm for human-powered solving of large scale problems in domains such as image classification, data entry, optical character recognition, recommendation, and proofreading. Because these low-paid workers can be unreliable, nearly all crowdsourcers must devise schemes to increase confidence in their answers, typically by assigning each task multiple times and combining the answers in some way such as majority voting. In this paper, we consider a general model of such crowdsourcing tasks, and pose the problem of minimizing the total price (i.e., number of task assignments) that must be paid to achieve a target overall reliability. We give a new algorithm for deciding which tasks to assign to which workers and for inferring correct answers from the workers' answers. We show that our algorithm significantly outperforms majority voting and, in fact, is asymptotically optimal through comparison to an oracle that knows the reliability of every worker.