Crowdsourced PAC Learning under Classification Noise

Crowdsourced PAC Learning under Classification Noise
复制标题

DOI:
10.1609/hcomp.v7i1.5279
复制
发表时间:
2019-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Shelby Heinecke;L. Reyzin
Shelby Heinecke;L. Reyzin
中科院分区:
其他
文献类型:
--
作者:
Shelby Heinecke;L. Reyzin

文献摘要

被引文献

相似文献

在本文中,我们从众包产生的标签中分析了PAC可学习性。在我们的环境中,未标记的示例是从分布中绘制的,标签是由在分类噪声下操作的工人众包,每个人都有自己的噪声参数。我们开发了一种端到端的众包PAC学习算法,该算法将未标记的数据点作为输入,并输出训练有素的分类器。我们的三步算法结合了多数投票,纯探索土匪和嘈杂的PAC学习。我们证明,在这种情况下,工人在PAC学习中标记的任务数量有几种保证,并表明我们的算法通过减少给工人的任务总数来改善基线。我们通过探索其在其他现实的众包环境中的应用来证明我们的算法的鲁棒性。
In this paper, we analyze PAC learnability from labels produced by crowdsourcing. In our setting, unlabeled examples are drawn from a distribution and labels are crowdsourced from workers who operate under classification noise, each with their own noise parameter. We develop an end-to-end crowdsourced PAC learning algorithm that takes unlabeled data points as input and outputs a trained classifier. Our three-step algorithm incorporates majority voting, pure-exploration bandits, and noisy-PAC learning. We prove several guarantees on the number of tasks labeled by workers for PAC learning in this setting and show that our algorithm improves upon the baseline by reducing the total number of tasks given to workers. We demonstrate the robustness of our algorithm by exploring its application to additional realistic crowdsourcing settings.