Object-Aware Cropping for Self-Supervised Learning

Object-Aware Cropping for Self-Supervised Learning
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Shlok Kumar Mishra;Anshul B. Shah;Ankan Bansal;Abhyuday N. Jagannatha;Abhishek Sharma;David Jacobs;Dilip Krishnan
Shlok Kumar Mishra;Anshul B. Shah;Ankan Bansal;Abhyuday N. Jagannatha;Abhishek Sharma;David Jacobs;Dilip Krishnan
中科院分区:
其他
文献类型:
--
作者:
Shlok Kumar Mishra;Anshul B. Shah;Ankan Bansal;Abhyuday N. Jagannatha;Abhishek Sharma;David Jacobs;Dilip Krishnan

文献摘要

相似文献

最近自我监督学习成功的一个核心组成部分是裁剪数据增强,它选择图像的子区域作为自我监督损失中的积极视图。基本的假设是,给定图像的随机裁剪和调整大小的区域共享关于感兴趣对象的信息,学习的表示将捕获这些信息。这一假设在像ImageNet这样的数据集中得到了最大程度的满足,其中有一个大的、居中的对象,它很可能出现在完整图像的随机裁剪中。然而,在诸如OpenImages或COCO等更能代表真实世界未整理数据的其他数据集中,图像中通常存在多个小对象。在这项工作中,我们表明,基于通常的随机裁剪的自我监督学习在这样的数据集上表现不佳。我们建议用从目标建议算法获得的作物来代替随机作物中的一种或两种。这鼓励模型学习对象和场景级别的语义表示。使用这种方法,我们称之为对象感知裁剪,在分类和对象检测基准上比场景裁剪有显著的改进。例如,在OpenImages上,与基于MoCo-v2的预训练的随机场景级裁剪相比,我们的方法获得了8.8%的MAP改进。与最先进的自监督学习方法相比,我们还显示出在COCO和PASCAL-VOC目标检测和分割任务上的显著改进。我们的方法是高效、简单和通用的,可以用于大多数现有的对比性和非对比性自我监督学习框架。
A core component of the recent success of self-supervised learning is cropping data augmentation, which selects sub-regions of an image to be used as positive views in the self-supervised loss. The underlying assumption is that randomly cropped and resized regions of a given image share information about the objects of interest, which the learned representation will capture. This assumption is mostly satisfied in datasets such as ImageNet where there is a large, centered object, which is highly likely to be present in random crops of the full image. However, in other datasets such as OpenImages or COCO, which are more representative of real world uncurated data, there are typically multiple small objects in an image. In this work, we show that self-supervised learning based on the usual random cropping performs poorly on such datasets. We propose replacing one or both of the random crops with crops obtained from an object proposal algorithm. This encourages the model to learn both object and scene level semantic representations. Using this approach, which we call object-aware cropping, results in significant improvements over scene cropping on classification and object detection benchmarks. For example, on OpenImages, our approach achieves an improvement of 8.8% mAP over random scene-level cropping using MoCo-v2 based pre-training. We also show significant improvements on COCO and PASCAL-VOC object detection and segmentation tasks over the state-of-the-art self-supervised learning approaches. Our approach is efficient, simple and general, and can be used in most existing contrastive and non-contrastive self-supervised learning frameworks.