D2C: Diffusion-Decoding Models for Few-Shot Conditional Generation

D2C: Diffusion-Decoding Models for Few-Shot Conditional Generation
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Abhishek Sinha;Jiaming Song;Chenlin Meng;Stefano Ermon
Abhishek Sinha;Jiaming Song;Chenlin Meng;Stefano Ermon
中科院分区:
其他
文献类型:
--
作者:
Abhishek Sinha;Jiaming Song;Chenlin Meng;Stefano Ermon

文献摘要

被引文献

相似文献

高维图像的条件生成模型有许多应用,但是从条件到图像的监督信号的获取可能是昂贵的。本文描述了具有对比表示的扩散解码模型(D2C),这是一种用于训练无条件变分自编码器(VAE)以生成少量条件图像的范例。D2C在潜在表示上使用学习的基于扩散的先验来改进生成,并使用对比自监督学习来提高表示质量。D2C可以通过从最少100个带标签的示例中学习来适应以标签或操作约束为条件的新生成任务。在从新标签的条件生成上,D2C实现了比最先进的VAE和扩散模型更优越的上级性能。在条件图像处理方面,D2C生成比StyleGAN 2生成快两个数量级,并且在双盲研究中受到50% - 60%的人类评估者的青睐。
Conditional generative models of high-dimensional images have many applications, but supervision signals from conditions to images can be expensive to acquire. This paper describes Diffusion-Decoding models with Contrastive representations (D2C), a paradigm for training unconditional variational autoencoders (VAEs) for few-shot conditional image generation. D2C uses a learned diffusion-based prior over the latent representations to improve generation and contrastive self-supervised learning to improve representation quality. D2C can adapt to novel generation tasks conditioned on labels or manipulation constraints, by learning from as few as 100 labeled examples. On conditional generation from new labels, D2C achieves superior performance over state-of-the-art VAEs and diffusion models. On conditional image manipulation, D2C generations are two orders of magnitude faster to produce over StyleGAN2 ones and are preferred by 50% - 60% of the human evaluators in a double-blind study.