Learning Rare Category Classifiers on a Tight Labeling Budget

Learning Rare Category Classifiers on a Tight Labeling Budget
复制标题

DOI:
10.1109/iccv48922.2021.00831
复制
发表时间:
2021-10
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Ravi Teja Mullapudi;Fait Poms;W. Mark;Deva Ramanan;Kayvon Fatahalian
Ravi Teja Mullapudi;Fait Poms;W. Mark;Deva Ramanan;Kayvon Fatahalian
中科院分区:
其他
文献类型:
--
作者:
Ravi Teja Mullapudi;Fait Poms;W. Mark;Deva Ramanan;Kayvon Fatahalian

文献摘要

相似文献

许多现实世界的机器学习部署都面临着用少量标签预算训练罕见类别模型的挑战。在这些设置中,通常可以访问大量未标记的数据,因此考虑半监督或主动学习方法来减少人工标记工作是很有吸引力的。然而,先前的方法做出了两个在实践中通常不成立的假设: (a) 可以访问适量的标记数据来引导学习,(b) 每张图像都属于一个共同感兴趣的类别。在本文中,我们考虑这样一种情况:我们从稀有类别的少至 5 个标记的正例数据和大量未标记数据(其中 99.9% 为负例)开始。我们提出了一种主动半监督方法,用于在这种具有挑战性的环境中构建准确的模型。我们的方法利用了两个关键思想:(a)在最有效的地方利用人力和机器的努力;人类标签用于识别“大海捞针”的阳性结果,而机器生成的伪标签用于识别阴性结果。 (b) 采用最近提出的表示学习技术来处理极其不平衡的人类标记数据,以迭代地训练带有噪声机器标记数据的模型。我们将我们的方法与之前的主动学习和半监督方法进行比较,证明每单位标记工作的准确性显着提高,特别是在标记预算紧张的情况下。
Many real-world ML deployments face the challenge of training a rare category model with a small labeling budget. In these settings, there is often access to large amounts of unlabeled data, therefore it is attractive to consider semi-supervised or active learning approaches to reduce human labeling effort. However, prior approaches make two assumptions that do not often hold in practice; (a) one has access to a modest amount of labeled data to bootstrap learning and (b) every image belongs to a common category of interest. In this paper, we consider the scenario where we start with as-little-as five labeled positives of a rare category and a large amount of unlabeled data of which 99.9% of it is negatives. We propose an active semi-supervised method for building accurate models in this challenging setting. Our method leverages two key ideas: (a) Utilize human and machine effort where they are most effective; human labels are used to identify "needle-in-a-haystack" positives, while machine-generated pseudo-labels are used to identify negatives. (b) Adapt recently proposed representation learning techniques for handling extremely imbalanced human labeled data to iteratively train models with noisy machine labeled data. We compare our approach with prior active learning and semi-supervised approaches, demonstrating significant improvements in accuracy per unit labeling effort, particularly on a tight labeling budget.