Learning to Overcome Noise in Weak Caption Supervision for Object Detection
Learning to Overcome Noise in Weak Caption Supervision for Object Detection
复制标题
DOI:
10.1109/tpami.2022.3187350
复制
发表时间:
2022-06
影响因子:
23.6
通讯作者:
Mesut Erhan Unal;Keren Ye;Mingda Zhang;Christopher Thomas;Adriana Kovashka;Wei Li;Danfeng Qin;
中科院分区:
文献类型:
--
作者:
Mesut Erhan Unal;Keren Ye;Mingda Zhang;Christopher Thomas;Adriana Kovashka;Wei Li;Danfeng Qin;
We propose the first mechanism to train object detection models from weak supervision in the form of captions at the image level. Language-based supervision for detection is appealing and inexpensive: many blogs with images and descriptive text written by human users exist. However, there is significant noise in this supervision: captions do not mention all objects that are shown, and may mention extraneous concepts. We first propose a technique to determine which image-caption pairs provide suitable signal for supervision. We further propose several complementary mechanisms to extract image-level pseudo labels for training from the caption. Finally, we train an iterative weakly-supervised object detection model from these image-level pseudo labels. We use captions from four datasets (COCO, Flickr30K, MIRFlickr1M, and Conceptual Captions) whose level of noise varies. We evaluate our approach on two object detection datasets. Weighting the labels extracted from different captions provides a boost over treating all captions equally. Further, our primary proposed technique for inferring pseudo labels for training at the image level, outperforms alternative techniques under a wide variety of settings. Both techniques generalize to datasets beyond the one they were trained on.