Multi-Label Active Learning from Crowds

Multi-Label Active Learning from Crowds
复制标题

DOI:
--
复制
发表时间:
2015-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Shao-Yuan Li;Yuan Jiang;Zhi-Hua Zhou
Shao-Yuan Li;Yuan Jiang;Zhi-Hua Zhou
中科院分区:
其他
文献类型:
--
作者:
Shao-Yuan Li;Yuan Jiang;Zhi-Hua Zhou

文献摘要

被引文献

相似文献

多标签主动学习是通过从oracle中最优地选择最有价值的实例来查询其标签,从而降低标签成本的研究热点。在本文中,我们考虑了在众包环境下基于池的多标签主动学习,在主动查询过程中,可以使用多个具有不同专业知识的低成本不完美注释器来标记,而不是依靠高成本的oracle来获取基础事实。为了解决这个问题,我们提出了MAC (Multi-label Active learning from Crowds)方法,该方法结合了标签相关性的局部影响,在多标签分类器和注释器上建立了一个概率模型。基于这个模型,我们可以估计实例的标签以及每个注释者的专业知识。然后,我们提出了考虑实例不确定性/多样性和注释器可靠性的实例选择和注释器选择准则,从而在最有价值的实例中查询最可靠的注释器。实验结果证明了该方法的有效性。
Multi-label active learning is a hot topic in reducing the label cost by optimally choosing the most valuable instance to query its label from an oracle. In this paper, we consider the poolbased multi-label active learning under the crowdsourcing setting, where during the active query process, instead of resorting to a high cost oracle for the ground-truth, multiple low cost imperfect annotators with various expertise are available for labeling. To deal with this problem, we propose the MAC (Multi-label Active learning from Crowds) approach which incorporate the local influence of label correlations to build a probabilistic model over the multi-label classifier and annotators. Based on this model, we can estimate the labels for instances as well as the expertise of each annotator. Then we propose the instance selection and annotator selection criteria that consider the uncertainty/diversity of instances and the reliability of annotators, such that the most reliable annotator will be queried for the most valuable instances. Experimental results demonstrate the effectiveness of the proposed approach.