GroupViT: Semantic Segmentation Emerges from Text Supervision

GroupViT: Semantic Segmentation Emerges from Text Supervision
复制标题

DOI:
10.1109/cvpr52688.2022.01760
复制
发表时间:
2022-02
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Jiarui Xu;Shalini De Mello;Sifei Liu;Wonmin Byeon;Thomas Breuel;J. Kautz;X. Wang
Jiarui Xu;Shalini De Mello;Sifei Liu;Wonmin Byeon;Thomas Breuel;J. Kautz;X. Wang
中科院分区:
其他
文献类型:
--
作者:
Jiarui Xu;Shalini De Mello;Sifei Liu;Wonmin Byeon;Thomas Breuel;J. Kautz;X. Wang

文献摘要

被引文献

相似文献

图像处理和识别是视觉场景理解的重要组成部分,例如,用于对象检测和语义分割。对于端到端深度学习系统,图像区域的分组通常通过像素级识别标签的自上而下的监督隐式地发生。相反,在本文中,我们建议将分组机制带回到深度网络中,这允许语义段仅通过文本监督自动出现。我们提出了一个层次化的图像视觉Transformer(GroupViT),它超越了常规的网格结构表示,并学会将图像区域分组为逐渐变大的任意形状的片段。我们通过对比损失在大规模图像-文本数据集上与文本编码器联合训练GroupViT。只有文本监督,没有任何像素级注释,GroupViT学习将语义区域分组在一起,并以零射击的方式成功地转移到语义分割的任务,即,而无需任何进一步的微调。它在PASCAL VOC 2012上实现了52.3% mIoU的零射击精度,在PASCAL Context数据集上实现了22.4% mIoU,并且与需要更高级别监督的最先进的迁移学习方法相比具有竞争力。我们在https://github.com/NVlabs/GroupViT上开源我们的代码。
Grouping and recognition are important components of visual scene understanding, e.g., for object detection and semantic segmentation. With end-to-end deep learning systems, grouping of image regions usually happens implicitly via top-down supervision from pixel-level recognition labels. Instead, in this paper, we propose to bring back the grouping mechanism into deep networks, which allows semantic segments to emerge automatically with only text supervision. We propose a hierarchical Grouping Vision Transformer (GroupViT), which goes beyond the regular grid structure representation and learns to group image regions into progressively larger arbitrary-shaped segments. We train GroupViT jointly with a text encoder on a large-scale image-text dataset via contrastive losses. With only text supervision and without any pixel-level annotations, GroupViT learns to group together semantic regions and successfully transfers to the task of semantic segmentation in a zero-shot manner, i.e., without any further fine-tuning. It achieves a zero-shot accuracy of 52.3% mIoU on the PASCAL VOC 2012 and 22.4% mIoU on PASCAL Context datasets, and performs competitively to state-of-the-art transfer-learning methods requiring greater levels of supervision. We open-source our code at https://github.com/NVlabs/GroupViT.