Your Diffusion Model is Secretly a Zero-Shot Classifier

Your Diffusion Model is Secretly a Zero-Shot Classifier
复制标题

DOI:
10.1109/iccv51070.2023.00210
复制
发表时间:
2023-03
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Alexander C. Li;Mihir Prabhudesai;Shivam Duggal;Ellis L Brown;Deepak Pathak
Alexander C. Li;Mihir Prabhudesai;Shivam Duggal;Ellis L Brown;Deepak Pathak
中科院分区:
其他
文献类型:
--
作者:
Alexander C. Li;Mihir Prabhudesai;Shivam Duggal;Ellis L Brown;Deepak Pathak

文献摘要

被引文献

相似文献

最近的大规模文本对图像扩散模型的浪潮显着提高了我们基于文本的图像生成能力。这些模型可以为各种提示而产生逼真的图像,并具有令人印象深刻的构图概括能力。迄今为止,几乎所有用例仅集中在抽样上。但是,扩散模型还可以提供条件密度估计,这对于图像生成以外的任务很有用。在本文中,我们表明,可以利用大规模的文本到图像扩散模型(如稳定扩散)的密度估计值,以执行零拍的分类,而无需任何额外的培训。我们称之为扩散分类器的生成分类方法,在各种基准测试基准上取得了强劲的结果,并且优于从扩散模型中提取知识的替代方法。尽管在零射击识别任务上的生成和歧视方法之间仍然存在差距,但我们基于扩散的方法具有更强的多模式组成推理能力,而不是竞争歧视性方法。最后,我们使用扩散分类器来从在Imagenet上训练的类条件扩散模型中提取标准分类器。这些模型接近SOTA判别分类器的性能,并表现出强大的“有效鲁棒性”来转移分配。总体而言,我们的结果是迈向使用歧视模型的生成型下游任务的一步。我们的网站上的结果和可视化:xudfusion-classifier.github.io/
The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhibit impressive compositional generalization abilities. Almost all use cases thus far have solely focused on sampling; however, diffusion models can also provide conditional density estimates, which are useful for tasks beyond image generation. In this paper, we show that the density estimates from large-scale text-to-image diffusion models like Stable Diffusion can be leveraged to perform zero-shot classification without any additional training. Our generative approach to classification, which we call Diffusion Classifier, attains strong results on a variety of benchmarks and outperforms alternative methods of extracting knowledge from diffusion models. Although a gap remains between generative and discriminative approaches on zero-shot recognition tasks, our diffusion-based approach has stronger multimodal compositional reasoning abilities than competing discriminative approaches. Finally, we use Diffusion Classifier to extract standard classifiers from class-conditional diffusion models trained on ImageNet. These models approach the performance of SOTA discriminative classifiers and exhibit strong "effective robustness" to distribution shift. Overall, our results are a step toward using generative over discriminative models for downstream tasks. Results and visualizations on our website: diffusion-classifier.github.io/