StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2

StyleGAN-V: A Continuous Video Generator with the Price, Image Quality and Perks of StyleGAN2
复制标题

DOI:
10.1109/cvpr52688.2022.00361
复制
发表时间:
2021-12
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Ivan Skorokhodov;S. Tulyakov;Mohamed Elhoseiny
Ivan Skorokhodov;S. Tulyakov;Mohamed Elhoseiny
中科院分区:
其他
文献类型:
--
作者:
Ivan Skorokhodov;S. Tulyakov;Mohamed Elhoseiny

文献摘要

被引文献

相似文献

视频显示了连续的事件,但大多数--如果不是全部--视频合成框架在时间上对它们进行了离散处理。在这项工作中,我们考虑了视频应该是时间连续的信号,并扩展了神经表示的范例来构建连续时间视频生成器。为此,我们首先通过位置嵌入的镜头来设计连续的运动表示。然后,我们探讨了在非常稀疏的视频上进行训练的问题,并证明了一个好的生成器可以通过每个片段只使用2帧来学习。之后,我们对传统的图像+视频鉴别器进行了反思,设计了一种通过简单地拼接帧的特征来聚合时间信息的整体鉴别器。这降低了训练成本,并为生成器提供了更丰富的学习信号,使首次直接对10242个视频进行训练成为可能。我们的模型建立在StyleGAN2之上,在获得几乎相同的图像质量的情况下,以相同的分辨率进行训练的成本仅比StyleGAN2高5%。此外,我们的潜在空间具有类似的性质,使得我们的方法可以在时间上传播空间操作。我们可以以任意的高帧速率生成任意长的视频,而以前的工作甚至难以以固定的速率生成帧。我们的模型在四个Mod-ern2562和一个10242分辨率的视频合成基准上进行了测试。就纯粹的衡量标准而言,它的平均≈表现比最接近的亚军高出30%。项目网站:https://universome.github.io/stylegan-v.
Videos show continuous events, yet most - if not all - video synthesis frameworks treat them discretely in time. In this work, we think of videos of what they should be - time-continuous signals, and extend the paradigm of neural representations to build a continuous-time video generator. For this, we first design continuous motion representations through the lens of positional embeddings. Then, we explore the question of training on very sparse videos and demon-strate that a good generator can be learned by using as few as 2 frames per clip. After that, we rethink the traditional image + video discriminators pair and design a holistic dis-criminator that aggregates temporal information by simply concatenating frames' features. This decreases the training cost and provides richer learning signal to the generator, making it possible to train directly on 10242 videos for the first time. We build our model on top of StyleGAN2 and it is just ≈5% more expensive to train at the same resolution while achieving almost the same image quality. Moreover, our latent space features similar properties, enabling spa-tial manipulations that our method can propagate in time. We can generate arbitrarily long videos at arbitrary high frame rate, while prior work struggles to generate even 64 frames at a fixed rate. Our model is tested on four mod-ern 2562 and one 10242 -resolution video synthesis bench-marks. In terms of sheer metrics, it performs on average ≈30% better than the closest runner-up. Project website: https://universome.github.io/stylegan-v.