X2Face: A network for controlling face generation by using images, audio, and pose codes

X2Face: A network for controlling face generation by using images, audio, and pose codes
复制标题

DOI:
10.1007/978-3-030-01261-8_41
复制
发表时间:
2018-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Olivia Wiles;A. S. Koepke;Andrew Zisserman
Olivia Wiles;A. S. Koepke;Andrew Zisserman
中科院分区:
其他
文献类型:
--
作者:
Olivia Wiles;A. S. Koepke;Andrew Zisserman

文献摘要

被引文献

相似文献

本文的目标是一个神经网络模型,该模型使用另一张脸或模态(例如音频)来控制给定面孔的姿势和表情。该模型可用于轻量级、复杂的视频和图像编辑。我们作出以下三项贡献。首先,我们介绍了一个网络,X2 Face,它可以控制一个源面(由一个或多个帧指定),使用驱动帧中的另一个面来产生一个生成的帧,该帧具有源帧的身份,但在驱动帧中的脸的姿势和表情。其次,我们提出了一种使用大量视频数据来训练网络完全自监督的方法。第三,我们证明了生成过程可以由其他形式驱动,例如音频或姿势代码,而无需对网络进行任何进一步的训练。将用另一张脸驱动一张脸的生成结果与最先进的自监督/监督方法进行比较。我们表明,我们的方法比其他方法更强大,因为它对输入数据的假设更少。我们还展示了使用我们的框架进行视频人脸编辑的示例。
The objective of this paper is a neural network model that controls the pose and expression of a given face, using another face or modality (eg audio). This model can then be used for lightweight, sophisticated video and image editing. We make the following three contributions. First, we introduce a network, X2Face, that can control a source face (specified by one or more frames) using another face in a driving frame to produce a generated frame with the identity of the source frame but the pose and expression of the face in the driving frame. Second, we propose a method for training the network fully self-supervised using a large collection of video data. Third, we show that the generation process can be driven by other modalities, such as audio or pose codes, without any further training of the network. The generation results for driving a face with another face are compared to state-of-the-art self-supervised/supervised methods. We show that our approach is more robust than other methods, as it makes fewer assumptions about the input data. We also show examples of using our framework for video face editing.