Learning to Decompose Visual Features with Latent Textual Prompts

Learning to Decompose Visual Features with Latent Textual Prompts
复制标题

DOI:
10.48550/arxiv.2210.04287
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Feng Wang;Manling Li;Xudong Lin;Hairong Lv;A. Schwing;Heng Ji
Feng Wang;Manling Li;Xudong Lin;Hairong Lv;A. Schwing;Heng Ji
中科院分区:
其他
文献类型:
--
作者:
Feng Wang;Manling Li;Xudong Lin;Hairong Lv;A. Schwing;Heng Ji

文献摘要

相似文献

像CLIP这样的预训练视觉语言模型的最新进展在学习可转移的视觉表示方面表现出了巨大的潜力。尽管如此,对于下游推理,CLIP类模型存在以下问题:1)在基于检索的推理过程中,如果文本描述不准确,则准确性和鲁棒性会降低(零触发协议的挑战);或者2)打破了既定的视觉语言对齐(线性探测的挑战)。为了解决这些问题,我们提出了分解特征提取(DeFo)。DeFo利用灵活数量的可学习嵌入作为文本输入,同时保持视觉语言双模型架构,使模型能够在特征级文本提示的帮助下学习分解的视觉特征。我们进一步使用一个额外的线性层来执行分类,允许语言输入的可扩展大小。我们的实证研究表明,DeFo的视觉语言模型的改进意义。例如,DeFo在ResNet-50主干上的ImageNet上获得了73.2%的测试准确率,而无需调整视觉和语言编码器的任何预训练权重,比zero-shot CLIP高出15.0%,比最先进的视觉语言提示调整方法高出7.6%。
Recent advances in pre-training vision-language models like CLIP have shown great potential in learning transferable visual representations. Nonetheless, for downstream inference, CLIP-like models suffer from either 1) degraded accuracy and robustness in the case of inaccurate text descriptions during retrieval-based inference (the challenge for zero-shot protocol); or 2) breaking the well-established vision-language alignment (the challenge for linear probing). To address them, we propose Decomposed Feature Prompting (DeFo). DeFo leverages a flexible number of learnable embeddings as textual input while maintaining the vision-language dual-model architecture, which enables the model to learn decomposed visual features with the help of feature-level textual prompts. We further use an additional linear layer to perform classification, allowing a scalable size of language inputs. Our empirical study shows DeFo's significance in improving the vision-language models. For example, DeFo obtains 73.2% test accuracy on ImageNet with a ResNet-50 backbone without tuning any pretrained weights of both the vision and language encoder, outperforming zero-shot CLIP by a large margin of 15.0%, and outperforming state-of-the-art vision-language prompt tuning method by 7.6%.