Detecting Twenty-thousand Classes using Image-level Supervision

Detecting Twenty-thousand Classes using Image-level Supervision
复制标题

DOI:
10.1007/978-3-031-20077-9_21
复制
发表时间:
2022-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Xingyi Zhou;Rohit Girdhar;Armand Joulin;Phillip Krahenbuhl;Ishan Misra
Xingyi Zhou;Rohit Girdhar;Armand Joulin;Phillip Krahenbuhl;Ishan Misra
中科院分区:
其他
文献类型:
--
作者:
Xingyi Zhou;Rohit Girdhar;Armand Joulin;Phillip Krahenbuhl;Ishan Misra

文献摘要

相似文献

当前的目标检测器由于检测数据集的规模小而在词汇量上受到限制。另一方面,图像分类器需要更大的词汇表,因为它们的数据集更大,更容易收集。我们提出了Detic,它只是在图像分类数据上训练检测器的分类器,从而将检测器的词汇表扩展到数万个概念。与以前的工作不同,Detic不需要复杂的分配方案来根据模型预测将图像标签分配给盒子,这使得它更容易实现,并与一系列检测架构和主干兼容。我们的研究结果表明,Detic产生优秀的检测器,即使没有框注释的类。它在开放词汇表和长尾检测基准上都优于以前的工作。在开放词汇LVIS基准测试中,Detic为所有类提供了2.4 mAP的增益,为新类提供了8.3 mAP的增益。在标准的LVIS基准测试中,Detic在对所有类或仅对罕见类进行评估时获得41.7 mAP,因此在样本较少的情况下缩小了对象类别的性能差距。这是第一次,我们用ImageNet数据集的所有21000个类训练检测器,并证明它可以推广到新的数据集,而无需微调。代码可在https://github.com/facebookresearch/Detic上获得。
Current object detectors are limited in vocabulary size due to the small scale of detection datasets. Image classifiers, on the other hand, reason about much larger vocabularies, as their datasets are larger and easier to collect. We proposeDetic, which simply trains the classifiers of a detector on image classification data and thus expands the vocabulary of detectors to tens of thousands of concepts. Unlike prior work, Detic does not need complex assignment schemes to assign image labels to boxes based on model predictions, making it much easier to implement and compatible with a range of detection architectures and backbones. Our results show that Detic yields excellent detectors even for classes without box annotations. It outperforms prior work on both open-vocabulary and long-tail detection benchmarks. Detic provides a gain of 2.4 mAP for all classes and 8.3 mAP for novel classes on the open-vocabulary LVIS benchmark. On the standard LVIS benchmark, Detic obtains 41.7 mAP when evaluated on all classes, or only rare classes, hence closing the gap in performance for object categories with few samples. For the first time, we train a detector with all the twenty-one-thousand classes of the ImageNet dataset and show that it generalizes to new datasets without finetuning. Code is available at https://github.com/facebookresearch/Detic.