Hyperbolic Contrastive Learning for Visual Representations beyond Objects

Hyperbolic Contrastive Learning for Visual Representations beyond Objects
复制标题

DOI:
10.1109/cvpr52729.2023.00661
复制
发表时间:
2022-12
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Songwei Ge;Shlok Kumar Mishra;Simon Kornblith;Chun-Liang Li;David Jacobs
Songwei Ge;Shlok Kumar Mishra;Simon Kornblith;Chun-Liang Li;David Jacobs
中科院分区:
其他
文献类型:
--
作者:
Songwei Ge;Shlok Kumar Mishra;Simon Kornblith;Chun-Liang Li;David Jacobs

文献摘要

相似文献

虽然自监督/非监督方法导致了视觉表征学习的快速发展,但这些方法通常使用相同的镜头来处理对象和场景。在本文中,我们专注于学习保持对象和场景之间结构的表示。由于观察到视觉上相似的物体在表征空间中很接近,我们认为场景和物体应该遵循基于它们的构成性的层次结构。为了利用这种结构,我们提出了一个对比学习框架,其中欧几里得损失用于学习对象表示,而双曲损失用于鼓励场景的表示在双曲空间中靠近其组成对象的表示。这个新的双曲线目标通过优化其范数的大小来鼓励场景对象在表征中的夸张。我们表明,当在COCO和OpenImages数据集上进行预训练时,双曲线损失改善了多个数据集和任务中多个基线的下游性能,包括图像分类、目标检测和语义分割。我们还展示了学习表示的属性允许我们以零镜头的方式解决涉及场景和对象之间的交互的各种视觉任务。
Although self-/un-supervised methods have led to rapid progress in visual representation learning, these methods generally treat objects and scenes using the same lens. In this paper, we focus on learning representations for objects and scenes that preserve the structure among them. Motivated by the observation that visually similar objects are close in the representation space, we argue that the scenes and objects should instead follow a hierarchical structure based on their compositionality. To exploit such a structure, we propose a contrastive learning framework where a Euclidean loss is used to learn object representations and a hyperbolic loss is used to encourage representations of scenes to lie close to representations of their constituent objects in a hyperbolic space. This novel hyperbolic objective encourages the scene-object hypernymy among the representations by optimizing the magnitude of their norms. We show that when pretraining on the COCO and OpenImages datasets, the hyperbolic loss improves downstream performance of several baselines across multiple datasets and tasks, including image classification, object detection, and semantic segmentation. We also show that the properties of the learned representations allow us to solve various vision tasks that involve the interaction between scenes and objects in a zero-shot fashion.