Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators

Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators
复制标题

DOI:
10.1109/iccv51070.2023.01462
复制
发表时间:
2023-03
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Levon Khachatryan;A. Movsisyan;Vahram Tadevosyan;Roberto Henschel;Zhangyang Wang;Shant Navasardyan;Humphrey Shi
Levon Khachatryan;A. Movsisyan;Vahram Tadevosyan;Roberto Henschel;Zhangyang Wang;Shant Navasardyan;Humphrey Shi
中科院分区:
其他
文献类型:
--
作者:
Levon Khachatryan;A. Movsisyan;Vahram Tadevosyan;Roberto Henschel;Zhangyang Wang;Shant Navasardyan;Humphrey Shi

文献摘要

被引文献

相似文献

最近的文本到视频生成方法依赖于计算量大的训练,并且需要大规模的视频数据集。在本文中,我们引入了一种新的任务--零镜头文本到视频的生成,并通过利用现有的文本到图像合成方法(如稳定扩散)的能力,提出了一种低成本的方法(无需任何训练或优化),使其适用于视频领域。我们的主要改进包括:(I)用运动动力学丰富生成帧的潜在代码,以保持全局场景和背景时间的一致性;以及(Ii)使用第一帧上每一帧的新的跨帧注意力来重新编程帧级别的自我注意,以保持前景对象的上下文、外观和身份。实验表明,这导致了低开销,但高质量和显著一致的视频生成。此外,我们的方法不仅限于文本到视频的合成,还适用于其他任务,如有条件的和特定于内容的视频生成,以及视频指令-Pix2Pix,即指令制导的视频编辑。实验表明,尽管我们没有对额外的视频数据进行训练,但我们的方法的性能与最近的方法相当,有时甚至更好。我们的代码在以下网址公开提供:https://github.com/Picsart-AI-Research/Text2Video-Zero.
Recent text-to-video generation approaches rely on computationally heavy training and require large-scale video datasets. In this paper, we introduce a new task, zero-shot text-to-video generation, and propose a low-cost approach (without any training or optimization) by leveraging the power of existing text-to-image synthesis methods (e.g. Stable Diffusion), making them suitable for the video domain. Our key modifications include (i) enriching the latent codes of the generated frames with motion dynamics to keep the global scene and the background time consistent; and (ii) reprogramming frame-level self-attention using a new cross-frame attention of each frame on the first frame, to preserve the context, appearance, and identity of the foreground object. Experiments show that this leads to low overhead, yet high-quality and remarkably consistent video generation. Moreover, our approach is not limited to text-to-video synthesis but is also applicable to other tasks such as conditional and content-specialized video generation, and Video Instruct-Pix2Pix, i.e., instruction-guided video editing. As experiments show, our method performs comparably or sometimes better than recent approaches, despite not being trained on additional video data. Our code is publicly available at: https://github.com/Picsart-AI-Research/Text2Video-Zero.