Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP

Open-Vocabulary Semantic Segmentation with Mask-adapted CLIP
复制标题

DOI:
10.1109/cvpr52729.2023.00682
复制
发表时间:
2022-10
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Feng Liang;Bichen Wu;Xiaoliang Dai;Kunpeng Li;Yinan Zhao;Hang Zhang;Peizhao Zhang;Péter Vajda;D. Marculescu
Feng Liang;Bichen Wu;Xiaoliang Dai;Kunpeng Li;Yinan Zhao;Hang Zhang;Peizhao Zhang;Péter Vajda;D. Marculescu
中科院分区:
其他
文献类型:
--
作者:
Feng Liang;Bichen Wu;Xiaoliang Dai;Kunpeng Li;Yinan Zhao;Hang Zhang;Peizhao Zhang;Péter Vajda;D. Marculescu

文献摘要

相似文献

开放词汇语义分割的目的是根据文本描述将图像分割成语义区域,这些区域在训练过程中可能没有看到。最近的两阶段方法首先生成类不可知掩码建议,然后利用预先训练的视觉语言模型,例如,CLIP,用于对掩蔽区域进行分类。我们确定这种模式的性能瓶颈是预训练的CLIP模型,因为它在掩码图像上表现不佳。为了解决这个问题,我们建议微调CLIP上的一组掩蔽的图像区域及其相应的文本描述。我们通过挖掘现有的图像标题数据集(例如,COCOCaptions),使用CLIP将掩蔽的图像区域与图像标题中的名词进行匹配。与具有固定类别的更精确和手动注释的分割标签(例如,COCO-Stuff),我们发现我们的嘈杂但多样的数据集可以更好地保留CLIP的泛化能力。沿着整个模型的微调,我们使用一种我们称之为蒙版提示调整的方法来利用蒙版图像中的“空白”区域。实验结果表明,模板快速调整带来了显着的改善,而无需修改CLIP的任何权重,它可以进一步改善一个完全微调的模型。特别是,当在COCO上训练并在ADE 20 K-150上评估时,我们的最佳模型实现了29.6%的mIoU,比之前的最先进模型高出8.5%。2017年,开放词汇通才模型首次在没有特定数据集调整的情况下与监督专家模型的性能相匹配。
Open-vocabulary semantic segmentation aims to segment an image into semantic regions according to text descriptions, which may not have been seen during training. Recent two-stage methods first generate class-agnostic mask proposals and then leverage pre-trained vision-language models, e.g., CLIP, to classify masked regions. We identify the performance bottleneck of this paradigm to be the pre-trained CLIP model, since it does not perform well on masked images. To address this, we propose to finetune CLIP on a collection of masked image regions and their corresponding text descriptions. We collect training data by mining an existing image-caption dataset (e.g., COCO Captions), using CLIP to match masked image regions to nouns in the image captions. Compared with the more precise and manually annotated segmentation labels with fixed classes (e.g., COCO-Stuff), we find our noisy but diverse dataset can better retain CLIP's generalization ability. Along with finetuning the entire model, we utilize the “blank” areas in masked images using a method we dub mask prompt tuning. Experiments demonstrate mask prompt tuning brings significant improvement without modifying any weights of CLIP, and it can further improve a fully finetuned model. In particular, when trained on COCO and evaluated on ADE20K-150, our best model achieves 29.6% mIoU, which is +8.5% higher than the previous state-of-the-art. For the first time, open-vocabulary generalist models match the performance of supervised specialist models in 2017 without dataset specific adaptations.