Consensus algorithms for biased labeling in crowdsourcing

Consensus algorithms for biased labeling in crowdsourcing
复制标题

DOI:
10.1016/j.ins.2016.12.026
复制
发表时间:
2017-03
期刊:
Inf. Sci.
影响因子:
--
通讯作者:
Jing Zhang;Victor S. Sheng;Qianmu Li;Jian Wu;Xindong Wu
Jing Zhang;Victor S. Sheng;Qianmu Li;Jian Wu;Xindong Wu
中科院分区:
其他
文献类型:
--
作者:
Jing Zhang;Victor S. Sheng;Qianmu Li;Jian Wu;Xindong Wu

文献摘要

被引文献

相似文献

虽然通过众包系统标注对象时,非专家标注者往往会表现出偏见,这一观点已成为公认的外行观点,但缺乏足够的证据观察和系统的实证研究。本文首先分析了来自不同领域的八个真实世界的数据集,这些数据集的类标签是从众包系统中收集的。我们的分析表明,有偏见的标签是一个系统的二元分类的趋势;换句话说,对于大量的注释者,他们的标签质量的负面类(应该是大多数)显着大于那些积极的类(少数)。因此,本文在这些数据集上实证研究了四种现有的基于EM的共识算法DS,GLAD,RY和ZenCrowd的性能。我们的调查表明,所有这些国家的最先进的算法忽略了数据集的潜在偏见的特点,表现不佳,虽然他们的模型的复杂性的系统。为了解决有偏标记的处理问题,本文进一步提出了一种新的共识算法,即自适应加权多数投票(AWMV),基于两个类的标记质量之间的统计差异。AWMV利用每个示例的多个噪声标签集合中的正标签的频率来获得偏置率,然后将从偏置率导出的权重分配给负标签和正标签。比较结果表明,本文提出的AWMV算法具有最好的整体性能。最后,本文指出了一些潜在的相关课题,为未来的研究。
Although it has become an accepted lay view that when labeling objects through crowdsourcing systems, non-expert annotators often exhibit biases, this argument lacks sufficient evidential observation and systematic empirical study. This paper initially analyzes eight real-world datasets from different domains whose class labels were collected from crowdsourcing systems. Our analyses show that biased labeling is a systematic tendency for binary categorization; in other words, for a large number of annotators, their labeling qualities on the negative class (supposed to be the majority) are significantly greater than are those on the positive class (minority). Therefore, the paper empirically studies the performance of four existing EM-based consensus algorithms, DS, GLAD, RY, and ZenCrowd, on these datasets. Our investigation shows that all of these state-of-the-art algorithms ignore the potential bias characteristics of datasets and perform badly although they model the complexity of the systems. To address the issue of handling biased labeling, the paper further proposes a novel consensus algorithm, namely adaptive weighted majority voting (AWMV), based on the statistical difference between the labeling qualities of the two classes. AWMV utilizes the frequency of positive labels in the multiple noisy label set of each example to obtain a bias rate and then assigns weights derived from the bias rate to negative and positive labels. Comparison results among the five consensus algorithms (AWMV and the four existing) show that the proposed AWMV algorithm has the best overall performance. Finally, this paper notes some potential related topics for future study.