Towards Image-to-Video Translation: A Structure-Aware Approach via Multi-stage Generative Adversarial Networks

Towards Image-to-Video Translation: A Structure-Aware Approach via Multi-stage Generative Adversarial Networks
复制标题

DOI:
10.1007/s11263-020-01328-9
复制
发表时间:
2020-04
影响因子:
19.5
通讯作者:
Long Zhao;Xi Peng;Yu Tian;M. Kapadia;Dimitris N. Metaxas
Long Zhao;Xi Peng;Yu Tian;M. Kapadia;Dimitris N. Metaxas
中科院分区:
计算机科学2区
文献类型:
--
作者:
Long Zhao;Xi Peng;Yu Tian;M. Kapadia;Dimitris N. Metaxas

文献摘要

相似文献

在本文中,我们考虑了图像到视频的转换问题,其中一个或一组输入图像被转换成包含单个物体运动的输出视频。特别是,我们专注于预测高级结构的运动,如面部表情和人体姿势。最近的方法要么是由结构条件驱动的,要么是基于时间的。条件驱动的方法通常训练转换网络来生成基于预测结构序列的未来帧。另一方面,基于时间的方法表明,使用从大量训练数据中学习到的时间知识的3D卷积网络可以生成短的高质量运动。在这项工作中,我们结合了这两种方法的优点,并提出了一个两阶段生成框架,其中视频从结构序列中预测,然后通过时间信号进行细化。为了在预测阶段更有效地建模运动,我们训练具有密集连接的网络来学习当前帧和未来帧之间的残余运动,从而避免学习与运动无关的细节。为了保证精炼阶段的时间一致性,我们采用秩损失进行对抗性训练。我们对两个图像到视频的翻译任务:面部表情重定向和人体姿势预测进行了广泛的实验。在这两项任务上的卓越结果都证明了我们方法的有效性。
In this paper, we consider the problem of image-to-video translation, where one or a set of input images are translated into an output video which contains motions of a single object. Especially, we focus on predicting motions conditioned by high-level structures, such as facial expression and human pose. Recent approaches are either driven by structural conditions or temporal-based. Condition-driven approaches typically train transformation networks to generate future frames conditioned on the predicted structural sequence. Temporal-based approaches, on the other hand, have shown that short high-quality motions can be generated using 3D convolutional networks with temporal knowledge learned from massive training data. In this work, we combine the benefits of both approaches and propose a two-stage generative framework where videos are forecast from the structural sequence and then refined by temporal signals. To model motions more efficiently in the forecasting stage, we train networks with dense connections to learn residual motions between the current and future frames, which avoids learning motion-irrelevant details. To ensure temporal consistency in the refining stage, we adopt the ranking loss for adversarial training. We conduct extensive experiments on two image-to-video translation tasks: facial expression retargeting and human pose forecasting. Superior results over the state of the art on both tasks demonstrate the effectiveness of our approach.