Video-to-Video Synthesis

Video-to-Video Synthesis
复制标题

DOI:
--
复制
发表时间:
2018-08
期刊:
--
影响因子:
--
通讯作者:
Ting-Chun Wang;Ming-Yu Liu;Jun-Yan Zhu;Guilin Liu;Andrew Tao;J. Kautz;Bryan Catanzaro
Ting-Chun Wang;Ming-Yu Liu;Jun-Yan Zhu;Guilin Liu;Andrew Tao;J. Kautz;Bryan Catanzaro
中科院分区:
其他
文献类型:
--
作者:
Ting-Chun Wang;Ming-Yu Liu;Jun-Yan Zhu;Guilin Liu;Andrew Tao;J. Kautz;Bryan Catanzaro

文献摘要

被引文献

相似文献

我们研究视频到视频合成的问题,其目标是从输入源视频(例如,语义分割掩码序列)转换为精确地描绘源视频的内容的输出真实感视频。虽然它的图像对应,图像到图像的合成问题,是一个热门话题,视频到视频的合成问题是在文献中探讨较少。在不理解时间动态的情况下,直接将现有的图像合成方法应用于输入视频通常会导致低视觉质量的时间上不相干的视频。在本文中,我们提出了一种新的生成对抗学习框架下的视频到视频合成方法。通过精心设计的生成器和嵌入式架构,再加上时空对抗目标,我们在各种输入格式(包括分割掩码,草图和姿势)上实现了高分辨率,逼真,时间连贯的视频结果。在多个基准上的实验表明,我们的方法相比强基线的优势。特别是,我们的模型能够合成长达30秒的街景2K分辨率视频,这大大提高了视频合成的最新水平。最后,我们将我们的方法应用于未来的视频预测,优于几个最先进的竞争系统。
We study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.g., a sequence of semantic segmentation masks) to an output photorealistic video that precisely depicts the content of the source video. While its image counterpart, the image-to-image synthesis problem, is a popular topic, the video-to-video synthesis problem is less explored in the literature. Without understanding temporal dynamics, directly applying existing image synthesis approaches to an input video often results in temporally incoherent videos of low visual quality. In this paper, we propose a novel video-to-video synthesis approach under the generative adversarial learning framework. Through carefully-designed generator and discriminator architectures, coupled with a spatio-temporal adversarial objective, we achieve high-resolution, photorealistic, temporally coherent video results on a diverse set of input formats including segmentation masks, sketches, and poses. Experiments on multiple benchmarks show the advantage of our method compared to strong baselines. In particular, our model is capable of synthesizing 2K resolution videos of street scenes up to 30 seconds long, which significantly advances the state-of-the-art of video synthesis. Finally, we apply our approach to future video prediction, outperforming several state-of-the-art competing systems.