Multi-task Crowdsourcing via an Optimization Framework

Multi-task Crowdsourcing via an Optimization Framework
复制标题

通过优化框架的多任务众包

DOI:
10.1145/3310227
复制
发表时间:
2019
影响因子:
3.6
通讯作者:
He, Jingrui
He, Jingrui
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zhou, Yao;Ying, Lei;He, Jingrui

文献摘要

参考文献

被引文献

相似文献

前所未有的数据量催生了将人类洞察力与机器学习技术相结合的趋势,这有助于使用众包来有效和高效地获取标签信息。众包的一个关键挑战是不同的工人素质,这决定了这些工人提供的标签信息的准确性。受同一组任务通常由同一组工人标记的观察的启发,我们研究了他们在多个相关任务中的行为,并提出了一个从任务和工人双重异质性中学习的优化框架。该方法使用权值张量来表示工人在多个任务中的行为,并通过利用其结构化信息来寻求张量的最优解。然后,我们提出了一个迭代算法来解决优化问题,并分析其计算复杂度。为了推断一个例子的真实标签,我们基于估计的张量构造一个工人系综,其决策将使用一组熵权进行加权。我们还证明了最耗时的更新块的梯度相对于工人是可分离的,这导致了具有更快速度的随机算法。此外,我们扩展的学习框架,以适应多类设置。最后,我们在几个数据集上测试了我们的框架的性能,并证明了它优于最先进的技术。
The unprecedented amounts of data have catalyzed the trend of combining human insights with machine learning techniques, which facilitate the use of crowdsourcing to enlist label information both effectively and efficiently. One crucial challenge in crowdsourcing is the diverse worker quality, which determines the accuracy of the label information provided by such workers. Motivated by the observations that same set of tasks are typically labeled by the same set of workers, we studied their behaviors across multiple related tasks and proposed an optimization framework for learning from task and worker dual heterogeneity. The proposed method uses a weight tensor to represent the workers’ behaviors across multiple tasks, and seeks to find the optimal solution of the tensor by exploiting its structured information. Then, we propose an iterative algorithm to solve the optimization problem and analyze its computational complexity. To infer the true label of an example, we construct a worker ensemble based on the estimated tensor, whose decisions will be weighted using a set of entropy weight. We also prove that the gradient of the most time-consuming updating block is separable with respect to the workers, which leads to a randomized algorithm with faster speed. Moreover, we extend the learning framework to accommodate to the multi-class setting. Finally, we test the performance of our framework on several datasets, and demonstrate its superiority over state-of-the-art techniques.
DOI: 10.1145/3219819.3219968
发表时间: 2018-07
期刊: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
影响因子: --
作者:
Dawei Zhou;Jingrui He;Hongxia Yang;Wei Fan
通讯作者: Dawei Zhou;Jingrui He;Hongxia Yang;Wei Fan
DOI: 10.1145/3097983.3098015
发表时间: 2017-08
期刊: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining
影响因子: --
作者:
Dawei Zhou;Si Zhang;M. Yildirim;S. Alcorn;Hanghang Tong;H. Davulcu;Jingrui He
通讯作者: Dawei Zhou;Si Zhang;M. Yildirim;S. Alcorn;Hanghang Tong;H. Davulcu;Jingrui He
DOI: 10.1145/2339530.2339581
发表时间: 2012-08
期刊: --
影响因子: --
作者:
Yao Hu;Debing Zhang;Jun Liu;Jieping Ye;Xiaofei He
通讯作者: Yao Hu;Debing Zhang;Jun Liu;Jieping Ye;Xiaofei He
一种实用的任意场景中任意目标物体计数方法
DOI: --
发表时间: 2013
期刊: IEEE International Conference on Multimedia and Expo
影响因子: --
作者:
Yao Zhou;Jiebo Luo
通讯作者: Jiebo Luo
DOI: 10.1137/1.9781611975673.2
发表时间: 2019-01
期刊: --
影响因子: --
作者:
Lecheng Zheng;Yu Cheng;Jingrui He
通讯作者: Lecheng Zheng;Yu Cheng;Jingrui He