Background Splitting: Finding Rare Classes in a Sea of Background

Background Splitting: Finding Rare Classes in a Sea of Background
复制标题

DOI:
10.1109/cvpr46437.2021.00795
复制
发表时间:
2020-08
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Ravi Teja Mullapudi;Fait Poms;W. Mark;Deva Ramanan;Kayvon Fatahalian
Ravi Teja Mullapudi;Fait Poms;W. Mark;Deva Ramanan;Kayvon Fatahalian
中科院分区:
其他
文献类型:
--
作者:
Ravi Teja Mullapudi;Fait Poms;W. Mark;Deva Ramanan;Kayvon Fatahalian

文献摘要

被引文献

相似文献

我们专注于为少数极其罕见的类别训练深度图像分类模型的问题。在这种常见的现实场景中,几乎所有图像都属于数据集中的背景类别。我们发现,在不平衡数据集上进行训练的最先进方法并不能在这种情况下产生准确的深度模型。我们的解决方案是在训练期间将大的、视觉上不同的背景分成许多较小的、视觉上相似的类别。我们通过使用额外的辅助损失扩展图像分类模型来实现这个想法,该辅助损失学习模仿训练集上预先存在的分类模型的预测。辅助损失不需要额外的人类标签,并通过强制模型区分所有训练集示例的辅助类别(包括属于主要稀有类别分类任务的整体背景的辅助类别)来规范共享网络主干中的特征学习。为了评估我们的方法,我们提供了 iNaturalist 和 Places365 数据集的修改版本,其中在训练期间只有一小部分稀有类别标签可用(所有其他图像都标记为背景)。通过共同学习识别选定的稀有类别和辅助类别,我们的方法生成的模型在 98.30% 的数据为背景时比最先进的不平衡学习基线高出 8.3 mAP 点,在 99.98% 的数据为背景时比微调基线高出 42.3 mAP 点。
We focus on the problem of training deep image classification models for a small number of extremely rare categories. In this common, real-world scenario, almost all images belong to the background category in the dataset. We find that state-of-the-art approaches for training on imbalanced datasets do not produce accurate deep models in this regime. Our solution is to split the large, visually diverse background into many smaller, visually similar categories during training. We implement this idea by extending an image classification model with an additional auxiliary loss that learns to mimic the predictions of a pre-existing classification model on the training set. The auxiliary loss requires no additional human labels and regularizes feature learning in the shared network trunk by forcing the model to discriminate between auxiliary categories for all training set examples, including those belonging to the monolithic background of the main rare category classification task. To evaluate our method we contribute modified versions of the iNaturalist and Places365 datasets where only a small subset of rare category labels are available during training (all other images are labeled as background). By jointly learning to recognize both the selected rare categories and auxiliary categories, our approach yields models that perform 8.3 mAP points higher than state-of-the-art imbalanced learning baselines when 98.30% of the data is background, and up to 42.3 mAP points higher than fine-tuning baselines when 99.98% of the data is background.