PostMe: Unsupervised Dynamic Microtask Posting For Efficient and Reliable Crowdsourcing

PostMe: Unsupervised Dynamic Microtask Posting For Efficient and Reliable Crowdsourcing
复制标题

PostMe:无监督动态微任务发布,实现高效可靠的众包

DOI:
10.1109/bigdata55660.2022.10020590
复制
发表时间:
2022
期刊:
Proc. The 2022 IEEE International Conference on Big Data Workshop HMData 2022
影响因子:
--
通讯作者:
Tetsuji Ogawa
Tetsuji Ogawa
中科院分区:
--
文献类型:
--
作者:
Ryo Yanagisawa;Susumu Saito;Teppei Nakano;Tetsunori Kobayashi;Tetsuji Ogawa

文献摘要

相似文献

即使经过十多年的众包研究,我们也没有一个众包数据标注低成本质量保证的标准框架。本文提出了一种动态微任务发布的无监督学习方法,该方法允许每个微任务根据数据难度调整自己收集的响应数。由于众包数据标签可能包含错误,研究人员通常采用多数投票的方式,汇总多名工作人员的回答来计算最终的标签。然而,这种技术涉及到标签准确性和成本之间的权衡。本文提出了一种动态微任务发布模型,该模型在保持标注准确性的同时减少了收集到的响应总数;我们还旨在通过“无监督”方法获得模型,该方法不需要通过微任务发布的经验来训练标记有基础事实的数据。我们对牲畜监测图像的注释模拟表明,我们的方法取得了i)与需要使用标记数据进行模型训练的监督方法相当的学习性能,以及ii)与简单多数投票相比,在不降低准确性的情况下显著降低了成本。
Even after over a decade of many crowdsourcing researches, we have no standard framework for low-cost quality assurance in crowdsourced data annotation. This paper proposes an unsupervised learning method for dynamic microtask posting which allows each microtask to adjust their own number of collected responses based on the data difficulty. Since crowdsourced data labels are likely to contain errors, researchers often employ majority voting that aggregates responses from multiple workers to calculate a final l abel. T his t echnique, h owever, i nvolves a trade-off between label accuracy and cost. This paper presents a dynamic microtask posting model that reduces the total number of collected responses while maintaining the labeling accuracy; we also aim to obtain the model with an “unsupervised” approach, which does not require training through experience of microtask posting for data labeled with ground-truths. Our simulation in annotating livestock surveillance images demonstrated that our approach achieved i) comparable learning performance to that of the supervised approach that required model training with labeled data, and ii) a significant c ost r eduction without degrading accuracy in comparison to simple majority voting.