Realistic Speech-Driven Facial Animation with GANs

Realistic Speech-Driven Facial Animation with GANs
复制标题

DOI:
10.1007/s11263-019-01251-8
复制
发表时间:
2019-06
影响因子:
19.5
通讯作者:
Konstantinos Vougioukas;Stavros Petridis;M. Pantic
Konstantinos Vougioukas;Stavros Petridis;M. Pantic
中科院分区:
计算机科学2区
文献类型:
--
作者:
Konstantinos Vougioukas;Stavros Petridis;M. Pantic

文献摘要

被引文献

相似文献

语音驱动的人脸动画是基于语音信号自动合成说话人物的过程。这个领域的大部分工作创建了从音频特征到视觉特征的映射。这种方法通常需要使用计算机图形技术进行后处理,以产生逼真的结果,尽管依赖于对象。我们提出了一个端到端的系统,生成一个会说话的头的视频,只使用一个人的静态图像和包含语音的音频剪辑,而不依赖于手工制作的中间功能。我们的方法生成具有(a)与音频同步的嘴唇运动和(B)自然面部表情(例如眨眼和眉毛运动)的视频。我们的时间GAN使用3个鉴别器,专注于实现详细的帧,视听同步和逼真的表情。我们使用消融研究量化了模型中每个组件的贡献,并提供了对模型潜在表示的见解。生成的视频根据清晰度、重建质量、唇读准确度、同步性以及生成自然眨眼的能力进行评估。
Speech-driven facial animation is the process that automatically synthesizes talking characters based on speech signals. The majority of work in this domain creates a mapping from audio features to visual features. This approach often requires post-processing using computer graphics techniques to produce realistic albeit subject dependent results. We present an end-to-end system that generates videos of a talking head, using only a still image of a person and an audio clip containing speech, without relying on handcrafted intermediate features. Our method generates videos which have (a) lip movements that are in sync with the audio and (b) natural facial expressions such as blinks and eyebrow movements. Our temporal GAN uses 3 discriminators focused on achieving detailed frames, audio-visual synchronization, and realistic expressions. We quantify the contribution of each component in our model using an ablation study and we provide insights into the latent representation of the model. The generated videos are evaluated based on sharpness, reconstruction quality, lip-reading accuracy, synchronization as well as their ability to generate natural blinks.