A Convex Formulation for Learning from Crowds

A Convex Formulation for Learning from Crowds
复制标题

DOI:
10.1609/aaai.v26i1.8105
复制
发表时间:
2012-07
期刊:
Proceedings of the AAAI Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Hiroshi Kajino;Yuta Tsuboi;H. Kashima
Hiroshi Kajino;Yuta Tsuboi;H. Kashima
中科院分区:
其他
文献类型:
--
作者:
Hiroshi Kajino;Yuta Tsuboi;H. Kashima

文献摘要

被引文献

相似文献

最近,众包服务经常被用来收集大量的标记数据用于机器学习,因为它们为我们提供了一种简单的方法,可以在很短的时间内以很低的成本获得标签。众包的使用给机器学习带来了新的挑战,即如何应对众包生成的数据质量参差不齐的问题。虽然最近有许多尝试来解决多个工人的质量问题,只有少数现有的方法考虑直接从这样的噪声数据学习分类器的问题。所有这些方法都将真实标签建模为潜在变量,这导致了非凸优化问题。在本文中,我们提出了一个凸优化配方,从人群中学习,而不估计真正的标签,通过引入个人模型的个人人群工作者。我们还设计了一个有效的迭代方法来解决凸优化问题,利用多分类器中的条件独立结构。我们评估所提出的方法对三个竞争的方法在合成数据集和一个真实的众包数据集,并证明所提出的方法优于其他三种方法。
Recently crowdsourcing services are often used to collect a large amount of labeled data for machine learning, since they provide us an easy way to get labels at very low cost and in a short period. The use of crowdsourcing has introduced a new challenge in machine learning, that is, coping with the variable quality of crowd-generated data. Although there have been many recent attempts to address the quality problem of multiple workers, only a few of the existing methods consider the problem of learning classifiers directly from such noisy data. All these methods modeled the true labels as latent variables, which resulted in non-convex optimization problems. In this paper, we propose a convex optimization formulation for learning from crowds without estimating the true labels by introducing personal models of the individual crowd workers. We also devise an efficient iterative method for solving the convex optimization problems by exploiting conditional independence structures in multiple classifiers. We evaluate the proposed method against three competing methods on synthetic data sets and a real crowdsourced data set and demonstrate that the proposed method outperforms the other three methods.