Visual Classification via Description from Large Language Models

Visual Classification via Description from Large Language Models
复制标题

通过大型语言模型的描述进行视觉分类

DOI:
--
复制
发表时间:
2022
期刊:
International Conference on Learning Representations
影响因子:
--
通讯作者:
Carl Vondrick
Carl Vondrick
中科院分区:
--
文献类型:
--
作者:
Sachit Menon;Carl Vondrick

文献摘要

参考文献

被引文献

相似文献

视觉语言模型(VLM)(如CLIP)在使用标准零拍摄分类过程的各种识别任务上表现出了良好的性能-计算查询图像与每个类别的嵌入词之间的相似性。由于只使用类别名称,他们忽视了利用语言提供的丰富的附加信息的上下文。这一程序对为什么选择某一类别没有中间的理解,而且也没有提供调整这一决定所使用的标准的机制。我们提出了一个替代框架的分类与VLMs,我们称之为分类的描述。我们要求VLM检查描述性特征,而不是广泛的类别:要找到老虎,请寻找它的条纹;它的爪子;等等。通过基于这些描述符的决策,我们可以提供额外的线索,鼓励使用我们想要使用的功能。在这个过程中,我们可以清楚地了解模型使用什么特征来构建其决策;它获得了某种程度的内在可解释性。我们查询大型语言模型(例如,GPT-3),以便以可扩展的方式获得这些描述符。大量的实验表明,我们的框架有许多优点过去的可解释性。我们展示了ImageNet在分布变化中的准确性提高;展示了调整VLM以识别训练过程中看不到的概念的能力;并说明了如何编辑描述符以有效减轻与基线相比的偏差。
Vision-language models (VLMs) such as CLIP have shown promising performance on a variety of recognition tasks using the standard zero-shot classification procedure -- computing similarity between the query image and the embedded words for each category. By only using the category name, they neglect to make use of the rich context of additional information that language affords. The procedure gives no intermediate understanding of why a category is chosen, and furthermore provides no mechanism for adjusting the criteria used towards this decision. We present an alternative framework for classification with VLMs, which we call classification by description. We ask VLMs to check for descriptive features rather than broad categories: to find a tiger, look for its stripes; its claws; and more. By basing decisions on these descriptors, we can provide additional cues that encourage using the features we want to be used. In the process, we can get a clear idea of what features the model uses to construct its decision; it gains some level of inherent explainability. We query large language models (e.g., GPT-3) for these descriptors to obtain them in a scalable way. Extensive experiments show our framework has numerous advantages past interpretability. We show improvements in accuracy on ImageNet across distribution shifts; demonstrate the ability to adapt VLMs to recognize concepts unseen during training; and illustrate how descriptors can be edited to effectively mitigate bias compared to the baseline.
DOI: 10.1109/cvpr52688.2022.00780
发表时间: 2021-09
期刊: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子: --
作者:
Mitchell Wortsman;Gabriel Ilharco;Mike Li;Jong Wook Kim;Hannaneh Hajishirzi;Ali Farhadi;Hongseok Namkoong-H
通讯作者: Mitchell Wortsman;Gabriel Ilharco;Mike Li;Jong Wook Kim;Hannaneh Hajishirzi;Ali Farhadi;Hongseok Namkoong-H