How to filter out random clickers in a crowdsourcing-based study?

How to filter out random clickers in a crowdsourcing-based study?
复制标题

如何在基于众包的研究中过滤掉随机点击者?

DOI:
10.1145/2442576.2442591
复制
发表时间:
2012
期刊:
Proceedings of the Royal Society of London. Series B: Biological Sciences
影响因子:
--
通讯作者:
Ji Soo Yi
Ji Soo Yi
中科院分区:
--
文献类型:
--
作者:
Sung;Hyokun Yun;Ji Soo Yi

文献摘要

被引文献

相似文献

基于众包的用户研究在信息可视化(InfoVis)和可视化分析(VA)中越来越受欢迎。然而,仍然不清楚如何应对一些不良的众包工作者,尤其是那些仅仅为了获取报酬而提交随机回答的人(以下简称随机点击者)。为了减轻随机点击者的影响,一些研究只是简单地排除异常值,但这种方法存在潜在风险,可能会丢失那些即使忠实参与但表现极端的参与者的数据。在本文中,我们评估了众包工作者回答的随机性程度,以推断该工作者是否是随机点击者。这样,我们就能够可靠地筛选出随机点击者,并发现基于众包的用户研究所得数据与受控实验室研究的数据具有可比性。我们还在一个众包平台上对1500名众包工作者测试了三种具有代表性的奖励方案(计件工资、定额和惩罚方案)以及四种不同的报酬水平(0.00美元、0.20美元、1.00美元和4.00美元),以研究不同支付条件对随机点击者数量的影响。结果表明,较高的报酬会降低随机点击者的比例,但这种参与质量的提高并不能证明相关额外成本的合理性。文中还详细讨论了如何优化支付方案和金额,以便经济地获取高质量数据。
Crowdsourcing-based user studies have become increasingly popular in information visualization (InfoVis) and visual analytics (VA). However, it is still unclear how to deal with some undesired crowdsourcing workers, especially those who submit random responses simply to gain wages (random clickers, henceforth). In order to mitigate the impacts of random clickers, several studies simply exclude outliers, but this approach has a potential risk of losing data from participants whose performances are extreme even though they participated faithfully. In this paper, we evaluated the degree of randomness in responses from a crowdsourcing worker to infer whether the worker is a random clicker. Thus, we could reliably filter out random clickers and found that resulting data from crowdsourcing-based user studies were comparable with those of a controlled lab study. We also tested three representative reward schemes (piece-rate, quota, and punishment schemes) with four different levels of compensations ($0.00, $0.20, $1.00, and $4.00) on a crowdsourcing platform with a total of 1,500 crowdsourcing workers to investigate the influences that different payment conditions have on the number of random clickers. The results show that higher compensations decrease the proportion of random clickers, but such increase in participation quality cannot justify the associated additional costs. A detailed discussion on how to optimize the payment scheme and amount to obtain high-quality data economically is provided.