Learning to Prompt CLIP for Monocular Depth Estimation: Exploring the Limits of Human Language

Learning to Prompt CLIP for Monocular Depth Estimation: Exploring the Limits of Human Language
复制标题

DOI:
10.1109/iccvw60793.2023.00218
复制
发表时间:
2023-10
期刊:
2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)
影响因子:
--
通讯作者:
Dylan Auty;K. Mikolajczyk
Dylan Auty;K. Mikolajczyk
中科院分区:
其他
文献类型:
--
作者:
Dylan Auty;K. Mikolajczyk

文献摘要

相似文献

CLIP是一个重要的视觉和语言训练框架,它对世界有着令人惊讶的普遍理解,在许多开放式任务中表现良好,很少或没有额外的训练。最近的技术已经使用CLIP通过使用深度相关提示来执行0次单目深度估计(MDE),但是在这些提示中使用人类语言呈现了不必要的人类偏差。在这项工作中,我们使用连续的可学习标记代替离散的人类语言单词来阐明这个问题。我们实现了显着的性能提升,并发现学习的令牌不整齐地映射到深度相关的人类语言,这意味着CLIP的深度概念是不是简洁地表达在人类语言。我们认为,这可能会扩展到其他CLIP概念,并相信这一发现将引发进一步的研究在所有开放式场景解释任务的非语言标记的使用和解释。代码可在https://github.com/DylanAuty/PromptLearningCLIP-MDE上获得
CLIP is a significant vision-and-language training framework that has shown surprisingly general understanding of the world, with good performance in many openended tasks with little or no additional training. A recent technique has used CLIP to perform 0-shot Monocular Depth Estimation (MDE) by using depth-related prompts, but the use of human language in these prompts presents an unnecessary human bias. In this work, we use continuous learnable tokens in place of discrete human-language words to shed light on the problem. We achieve a significant boost in performance, and find that the learned tokens do not map neatly to depth-related human language, implying that CLIP’s concept of depth is not succinctly expressible in human language. We posit that this may extend to other CLIP concepts, and believe that this finding will spark further research into both the use and interpretation of non-linguistic tokens in all open-ended scene interpretation tasks. Code is available at https://github.com/DylanAuty/PromptLearningCLIP-MDE