Adapative algorithms for crowd-aided categorization

Adapative algorithms for crowd-aided categorization
复制标题

人群辅助分类的自适应算法

DOI:
10.1007/s00778-021-00685-2
复制
发表时间:
--
期刊:
影响因子:
4.2
通讯作者:
Feng Jianhua
Feng Jianhua
中科院分区:
计算机科学2区
文献类型:
--
作者:
Li Yuanbing;Wu Xian;Jin Yifei;Li Jian;Li Guoliang;Feng Jianhua

文献摘要

相似文献

我们研究的问题,利用人类的智慧来分类大量的对象。在这个问题中,给定一个类别层次结构和一组对象,我们可以要求人类检查一个对象是否属于一个类别,我们的目标是找到最具成本效益的策略来为每个对象在层次结构中定位适当的类别,使得成本(即,要问人类的问题的数量)被最小化。这个问题有许多重要的应用,包括图像分类和产品分类。我们开发了一个在线框架,其中类别分布是逐步学习,从而自适应地确定一个有效的问题顺序。我们证明了,即使真正的类别分布是已知的,在计算上的问题是棘手的。我们开发了一个近似算法,并证明它实现了近似因子为2。我们还表明,有一个完全多项式时间的近似方案的问题。此外,我们提出了一个在线策略,实现了几乎相同的性能保证离线的最佳策略,即使事先没有知识的类别分布。我们开发有效的技术来容忍群体错误。在一个真实的众包平台上的实验证明了该方法的有效性。
We study the problem of utilizing human intelligence to categorize a large number of objects. In this problem, given a category hierarchy and a set of objects, we can ask humans to check whether an object belongs to a category, and our goal is to find the most cost-effective strategy to locate the appropriate category in the hierarchy for each object, such that the cost (i.e., the number of questions to ask humans) is minimized. There are many important applications of this problem, including image classification and product categorization. We develop an online framework, in which category distribution is gradually learned and thus an effective order of questions are adaptively determined. We prove that even if the true category distribution is known in advance, the problem is computationally intractable. We develop an approximation algorithm, and prove that it achieves an approximation factor of 2. We also show that there is a fully polynomial time approximation scheme for the problem. Furthermore, we propose an online strategy which achieves nearly the same performance guarantee as the offline optimal strategy, even if there is no knowledge about category distribution beforehand. We develop effective techniques to tolerate crowd errors. Experiments on a real crowdsourcing platform demonstrate the effectiveness of our method.