Emergent Correspondence from Image Diffusion

Emergent Correspondence from Image Diffusion
复制标题

DOI:
10.48550/arxiv.2306.03881
复制
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Luming Tang;Menglin Jia;Qianqian Wang;Cheng Perng Phoo;Bharath Hariharan
Luming Tang;Menglin Jia;Qianqian Wang;Cheng Perng Phoo;Bharath Hariharan
中科院分区:
其他
文献类型:
--
作者:
Luming Tang;Menglin Jia;Qianqian Wang;Cheng Perng Phoo;Bharath Hariharan

文献摘要

被引文献

相似文献

寻找图像之间的对应关系是计算机视觉中的一个基本问题。在这篇文章中,我们证明了在没有任何显式监督的情况下,图像扩散模型中出现了对应。我们提出了一种简单的策略来从扩散网络中提取这些隐含的知识作为图像特征,即扩散特征(DIFT),并使用它们来建立真实图像之间的对应关系。在不对特定于任务的数据或注释进行任何额外的微调或监督的情况下,DIFT能够在识别语义、几何和时间对应方面优于弱监督方法和竞争性现成特征。特别是在语义一致性方面,在具有挑战性的Sair-71K基准测试中,来自稳定扩散的DIFT能够分别比Dino和OpenCLIP高出19和14个精确点。它甚至在18个类别中的9个类别上超过了最先进的监督方法,而总体性能保持在同等水平。项目页面:https://diffusionfeatures.github.io
Finding correspondences between images is a fundamental problem in computer vision. In this paper, we show that correspondence emerges in image diffusion models without any explicit supervision. We propose a simple strategy to extract this implicit knowledge out of diffusion networks as image features, namely DIffusion FeaTures (DIFT), and use them to establish correspondences between real images. Without any additional fine-tuning or supervision on the task-specific data or annotations, DIFT is able to outperform both weakly-supervised methods and competitive off-the-shelf features in identifying semantic, geometric, and temporal correspondences. Particularly for semantic correspondence, DIFT from Stable Diffusion is able to outperform DINO and OpenCLIP by 19 and 14 accuracy points respectively on the challenging SPair-71k benchmark. It even outperforms the state-of-the-art supervised methods on 9 out of 18 categories while remaining on par for the overall performance. Project page: https://diffusionfeatures.github.io