Learning to Overcome Noise in Weak Caption Supervision for Object Detection

Learning to Overcome Noise in Weak Caption Supervision for Object Detection
复制标题

DOI:
10.1109/tpami.2022.3187350
复制
发表时间:
2022-06
影响因子:
23.6
通讯作者:
Mesut Erhan Unal;Keren Ye;Mingda Zhang;Christopher Thomas;Adriana Kovashka;Wei Li;Danfeng Qin;
Mesut Erhan Unal;Keren Ye;Mingda Zhang;Christopher Thomas;Adriana Kovashka;Wei Li;Danfeng Qin;
中科院分区:
计算机科学1区
文献类型:
--
作者:
Mesut Erhan Unal;Keren Ye;Mingda Zhang;Christopher Thomas;Adriana Kovashka;Wei Li;Danfeng Qin;

文献摘要

相似文献

我们提出了第一种机制,以图像级别的字幕形式训练弱监督下的目标检测模型。基于语言的检测监督很有吸引力,而且成本不高:存在许多由人类用户撰写的带有图像和描述性文字的博客。然而,在这种监督中有很大的噪音:标题没有提到所有显示的对象,可能会提到无关的概念。我们首先提出了一种技术来确定哪些图像-字幕对提供了合适的监督信号。我们进一步提出了几种互补的机制来从字幕中提取用于训练的图像级伪标签。最后,我们从这些图像级伪标签训练一个迭代的弱监督目标检测模型。我们使用四个数据集(COCO、Flickr30K、MIRFlickr1M和概念性字幕)中的字幕,它们的噪声级别各不相同。我们在两个目标检测数据集上对我们的方法进行了评估。与平等对待所有字幕相比,对从不同字幕中提取的标签进行加权可以提供更好的效果。此外,我们主要提出的用于在图像级别上推断用于训练的伪标签的技术在各种不同的设置下都优于其他技术。这两种技术都适用于超出其训练范围的数据集。
We propose the first mechanism to train object detection models from weak supervision in the form of captions at the image level. Language-based supervision for detection is appealing and inexpensive: many blogs with images and descriptive text written by human users exist. However, there is significant noise in this supervision: captions do not mention all objects that are shown, and may mention extraneous concepts. We first propose a technique to determine which image-caption pairs provide suitable signal for supervision. We further propose several complementary mechanisms to extract image-level pseudo labels for training from the caption. Finally, we train an iterative weakly-supervised object detection model from these image-level pseudo labels. We use captions from four datasets (COCO, Flickr30K, MIRFlickr1M, and Conceptual Captions) whose level of noise varies. We evaluate our approach on two object detection datasets. Weighting the labels extracted from different captions provides a boost over treating all captions equally. Further, our primary proposed technique for inferring pseudo labels for training at the image level, outperforms alternative techniques under a wide variety of settings. Both techniques generalize to datasets beyond the one they were trained on.