Learning to Prompt CLIP for Monocular Depth Estimation: Exploring the Limits of Human Language
Learning to Prompt CLIP for Monocular Depth Estimation: Exploring the Limits of Human Language
复制标题
DOI:
10.1109/iccvw60793.2023.00218
复制
发表时间:
2023-10
期刊:
影响因子:
--
通讯作者:
Dylan Auty;K. Mikolajczyk
中科院分区:
文献类型:
--
作者:
Dylan Auty;K. Mikolajczyk
CLIP is a significant vision-and-language training framework that has shown surprisingly general understanding of the world, with good performance in many openended tasks with little or no additional training. A recent technique has used CLIP to perform 0-shot Monocular Depth Estimation (MDE) by using depth-related prompts, but the use of human language in these prompts presents an unnecessary human bias. In this work, we use continuous learnable tokens in place of discrete human-language words to shed light on the problem. We achieve a significant boost in performance, and find that the learned tokens do not map neatly to depth-related human language, implying that CLIP’s concept of depth is not succinctly expressible in human language. We posit that this may extend to other CLIP concepts, and believe that this finding will spark further research into both the use and interpretation of non-linguistic tokens in all open-ended scene interpretation tasks. Code is available at https://github.com/DylanAuty/PromptLearningCLIP-MDE