Approximating the crowd

Approximating the crowd
复制标题

DOI:
10.1007/s10618-014-0354-1
复制
发表时间:
2014-09-01
影响因子:
4.8
通讯作者:
Hirsh, Haym
Hirsh, Haym
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ertekin, Seyda;Rudin, Cynthia;Hirsh, Haym

文献摘要

被引文献

相似文献

“接近人群”的问题是通过只询问其中的一个子集来估计人群的多数意见。接近人群的算法可以智能地将有限的预算用于众包任务。我们提出了一种名为“CrowdSense”的算法,它在一种在线方式下工作,一次来一个商品。CrowdSense根据探索/利用标准对人群的子集进行动态采样。该算法产生了接近大众意见的子集投票的加权组合。然后,我们介绍了CrowdSense的两个变体,它们进行了各种分布近似,以处理不同的人群特征。具体地说,第一个算法对大群体的标签者进行了统计独立近似,而第二个算法找到了当前子群体与群体多数投票一致的频率的下限。我们在CrowdSense和几条基线上的实验表明,通过从人群中具有代表性的子集收集意见,我们可以可靠地近似整个人群的投票。
The problem of "approximating the crowd" is that of estimating the crowd's majority opinion by querying only a subset of it. Algorithms that approximate the crowd can intelligently stretch a limited budget for a crowdsourcing task. We present an algorithm, "CrowdSense," that works in an online fashion where items come one at a time. CrowdSense dynamically samples subsets of the crowd based on an exploration/exploitation criterion. The algorithm produces a weighted combination of the subset's votes that approximates the crowd's opinion. We then introduce two variations of CrowdSense that make various distributional approximations to handle distinct crowd characteristics. In particular, the first algorithm makes a statistical independence approximation of the labelers for large crowds, whereas the second algorithm finds a lower bound on how often the current subcrowd agrees with the crowd's majority vote. Our experiments on CrowdSense and several baselines demonstrate that we can reliably approximate the entire crowd's vote by collecting opinions from a representative subset of the crowd.