Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation

Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation
复制标题

DOI:
10.1109/cvpr46437.2021.00416
复制
发表时间:
2021-04
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Hang Zhou;Yasheng Sun;Wayne Wu;Chen Change Loy;Xiaogang Wang;Ziwei Liu
Hang Zhou;Yasheng Sun;Wayne Wu;Chen Change Loy;Xiaogang Wang;Ziwei Liu
中科院分区:
其他
文献类型:
--
作者:
Hang Zhou;Yasheng Sun;Wayne Wu;Chen Change Loy;Xiaogang Wang;Ziwei Liu

文献摘要

被引文献

相似文献

虽然已经实现了精确的嘴唇同步任意主题的音频驱动的说话人脸生成,如何有效地驱动头部姿势的问题仍然存在。以前的方法依赖于预先估计的结构信息,如地标和3D参数,旨在生成个性化的节奏运动。然而,在极端条件下,这种估计信息的不准确性将导致退化问题。在本文中,我们提出了一个干净而有效的框架来生成姿态可控的说话脸。我们对非对齐的原始人脸图像进行操作,仅使用一张照片作为身份参考。关键是通过设计一个隐式的低维姿势代码来模块化视听表示。实质上,语音内容和头部姿态信息都位于联合非身份嵌入空间中。虽然语音内容信息可以通过学习视听模态之间的内在同步来定义,但我们发现姿势代码将在基于调制卷积的重建框架中补充学习。大量实验表明,我们的方法可以准确地生成口型同步的说话面孔,其姿势可由其他视频控制。此外,我们的模型具有多种高级功能,包括极端视图鲁棒性和说话面部前端化。
While accurate lip synchronization has been achieved for arbitrary-subject audio-driven talking face generation, the problem of how to efficiently drive the head pose remains. Previous methods rely on pre-estimated structural information such as landmarks and 3D parameters, aiming to generate personalized rhythmic movements. However, the inaccuracy of such estimated information under extreme conditions would lead to degradation problems. In this paper, we propose a clean yet effective framework to generate pose-controllable talking faces. We operate on non-aligned raw face images, using only a single photo as an identity reference. The key is to modularize audio-visual representations by devising an implicit low-dimension pose code. Substantially, both speech content and head pose information lie in a joint non-identity embedding space. While speech content information can be defined by learning the intrinsic synchronization between audio-visual modalities, we identify that a pose code will be complementarily learned in a modulated convolution-based reconstruction framework.Extensive experiments show that our method generates accurately lip-synced talking faces whose poses are controllable by other videos. Moreover, our model has multiple advanced capabilities including extreme view robustness and talking face frontalization.1