CLIP-Driven Prototype Network for Few-Shot Semantic Segmentation.

CLIP-Driven Prototype Network for Few-Shot Semantic Segmentation.
复制标题

DOI:
10.3390/e25091353
复制
发表时间:
2023-09-18
期刊:
Entropy (Basel, Switzerland)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

最近的研究表明,视觉文本预训练模型在传统的视觉任务中表现良好。CLIP作为最具影响力的研究成果,受到了研究者的广泛关注。由于其出色的视觉表现能力,许多最近的研究已经使用CLIP像素级的任务。我们探索CLIP在少镜头分割领域的潜在能力。目前主流的方法是利用支持和查询特征生成类原型,然后利用原型特征匹配图像特征。我们提出了一种新的方法,利用CLIP提取特定类的文本特征。然后,这些文本特征被用作训练样本来参与模型的训练过程。文本特征的加入使模型能够提取包含更丰富语义信息的特征,从而更容易捕捉潜在的类信息。为了更好地匹配查询图像特征,我们还提出了一种新的原型生成方法,在原型生成过程中,结合文本和图像的多模态融合特征。通过将来自图像的前景和背景信息与多模态支持原型相结合来生成自适应查询原型,从而允许更好地匹配图像特征并提高分割精度。我们提供了一个新的视角,在多模态场景中的少数镜头分割的任务。实验表明,我们提出的方法取得了良好的效果在两个常见的数据集,PASCAL-和COCO-。
Recent research has shown that visual–text pretrained models perform well in traditional vision tasks. CLIP, as the most influential work, has garnered significant attention from researchers. Thanks to its excellent visual representation capabilities, many recent studies have used CLIP for pixel-level tasks. We explore the potential abilities of CLIP in the field of few-shot segmentation. The current mainstream approach is to utilize support and query features to generate class prototypes and then use the prototype features to match image features. We propose a new method that utilizes CLIP to extract text features for a specific class. These text features are then used as training samples to participate in the model’s training process. The addition of text features enables model to extract features that contain richer semantic information, thus making it easier to capture potential class information. To better match the query image features, we also propose a new prototype generation method that incorporates multi-modal fusion features of text and images in the prototype generation process. Adaptive query prototypes were generated by combining foreground and background information from the images with the multi-modal support prototype, thereby allowing for a better matching of image features and improved segmentation accuracy. We provide a new perspective to the task of few-shot segmentation in multi-modal scenarios. Experiments demonstrate that our proposed method achieves excellent results on two common datasets, PASCAL- and COCO-.
DOI: 10.1109/tpami.2017.2699184
发表时间: 2018-04-01
影响因子: 23.6
作者:
Chen, Liang-Chieh;Papandreou, George;Yuille, Alan L.
通讯作者: Yuille, Alan L.
DOI: 10.1007/s11263-009-0275-4
发表时间: 2010-06-10
影响因子: 19.5
作者:
Everingham, Mark;Van Gool, Luc;Zisserman, Andrew
通讯作者: Zisserman, Andrew
DOI: 10.1109/tpami.2016.2644615
发表时间: 2017-12-01
影响因子: 23.6
作者:
Badrinarayanan, Vijay;Kendall, Alex;Cipolla, Roberto
通讯作者: Cipolla, Roberto
DOI: 10.1145/3065386
发表时间: 2017-06-01
影响因子: 22.7
作者:
Krizhevsky, Alex;Sutskever, Ilya;Hinton, Geoffrey E.
通讯作者: Hinton, Geoffrey E.