Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models

Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models
复制标题

DOI:
10.1109/iccv51070.2023.02096
复制
发表时间:
2023-05
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Songwei Ge;Seungjun Nah;Guilin Liu;Tyler Poon;Andrew Tao;Bryan Catanzaro;David Jacobs;Jia-Bin Huang;Ming-Yu Liu;Y. Balaji
Songwei Ge;Seungjun Nah;Guilin Liu;Tyler Poon;Andrew Tao;Bryan Catanzaro;David Jacobs;Jia-Bin Huang;Ming-Yu Liu;Y. Balaji
中科院分区:
其他
文献类型:
--
作者:
Songwei Ge;Seungjun Nah;Guilin Liu;Tyler Poon;Andrew Tao;Bryan Catanzaro;David Jacobs;Jia-Bin Huang;Ming-Yu Liu;Y. Balaji

文献摘要

相似文献

尽管在使用扩散模型生成高质量图像方面取得了巨大进展,但合成一系列既具有照片级真实感又具有时间连贯性的动画帧仍处于起步阶段。虽然用于图像生成的现成的十亿级数据集可用,但收集相同规模的类似视频数据仍然具有挑战性。此外,训练视频扩散模型在计算上比其图像对应物昂贵得多。在这项工作中,我们探索用视频数据微调预训练的图像扩散模型,作为视频合成任务的实用解决方案。我们发现,在视频扩散中,天真地将图像噪声扩展到视频噪声之前会导致次优性能。我们精心设计的视频噪声先验导致更好的性能。广泛的实验验证表明,我们的模型,保留自己的COrrelation(PYoCo),达到SOTA零拍摄文本到视频的结果,在UCF-101和MSR-VTT基准。它还在小规模UCF-101基准测试中实现了SOTA视频生成质量,使用比现有技术小10倍的模型,使用比现有技术少得多的计算。https://research.nvidia.com/labs/dir/pyoco/
Despite tremendous progress in generating high-quality images using diffusion models, synthesizing a sequence of animated frames that are both photorealistic and temporally coherent is still in its infancy. While off-the-shelf billion-scale datasets for image generation are available, collecting similar video data of the same scale is still challenging. Also, training a video diffusion model is computationally much more expensive than its image counterpart. In this work, we explore finetuning a pretrained image diffusion model with video data as a practical solution for the video synthesis task. We find that naively extending the image noise prior to video noise prior in video diffusion leads to sub-optimal performance. Our carefully designed video noise prior leads to substantially better performance. Extensive experimental validation shows that our model, Preserve Your Own COrrelation (PYoCo), attains SOTA zero-shot text-to-video results on the UCF-101 and MSR-VTT benchmarks. It also achieves SOTA video generation quality on the small-scale UCF-101 benchmark with a 10× smaller model using significantly less computation than the prior art. The project page is available at https://research.nvidia.com/labs/dir/pyoco/.