Versatile Diffusion: Text, Images and Variations All in One Diffusion Model

Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
复制标题

DOI:
10.1109/iccv51070.2023.00713
复制
发表时间:
2022-11
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Xingqian Xu;Zhangyang Wang;Eric Zhang;Kai Wang;Humphrey Shi
Xingqian Xu;Zhangyang Wang;Eric Zhang;Kai Wang;Humphrey Shi
中科院分区:
其他
文献类型:
--
作者:
Xingqian Xu;Zhangyang Wang;Eric Zhang;Kai Wang;Humphrey Shi

文献摘要

相似文献

扩散模型的最新进展在许多生成任务中树立了令人印象深刻的里程碑,DALL-E2,Imagen和Stable Diffusion等趋势作品引起了人们的极大兴趣。尽管环境变化迅速,但最近的新方法专注于扩展和性能而不是容量,因此需要单独的模型来执行单独的任务。在这项工作中,我们将现有的单流扩散管道扩展为多任务多模态网络,称为多功能扩散(VD),它可以在一个统一的模型中处理文本到图像,图像到文本和变化的多个流。VD的流水线设计实例化了一个统一的多流扩散框架,由可共享和可交换的层模块组成,使跨模式的通用性超越了图像和文本。通过大量的实验,我们证明了VD成功地实现了以下几点:a)VD优于基线方法,并以具有竞争力的质量处理其所有基本任务; B)VD实现了新的扩展,如风格和语义的解纠缠,双重和多上下文混合等; c)我们的多流多模态框架在图像和文本上的成功可能会激发进一步的基于扩散的通用AI研究。我们的代码和模型在https://github.com/SHI-Labs/Versatile-Diffusion上开源。
Recent advances in diffusion models have set an impressive milestone in many generation tasks, and trending works such as DALL-E2, Imagen, and Stable Diffusion have attracted great interest. Despite the rapid landscape changes, recent new approaches focus on extensions and performance rather than capacity, thus requiring separate models for separate tasks. In this work, we expand the existing single-flow diffusion pipeline into a multi-task multimodal network, dubbed Versatile Diffusion (VD), that handles multiple flows of text-to-image, image-to-text, and variations in one unified model. The pipeline design of VD instantiates a unified multi-flow diffusion framework, consisting of sharable and swappable layer modules that enable the crossmodal generality beyond images and text. Through extensive experiments, we demonstrate that VD successfully achieves the following: a) VD outperforms the baseline approaches and handles all its base tasks with competitive quality; b) VD enables novel extensions such as disentanglement of style and semantics, dual- and multi-context blending, etc.; c) The success of our multi-flow multimodal framework over images and text may inspire further diffusion-based universal AI research. Our code and models are open-sourced at https://github.com/SHI-Labs/Versatile-Diffusion.