Sound to Visual: Hierarchical Cross-Modal Talking Face Video Generation

Sound to Visual: Hierarchical Cross-Modal Talking Face Video Generation
复制标题

DOI:
--
复制
发表时间:
2019
期刊:
--
影响因子:
--
通讯作者:
Lele Chen;Haitian Zheng;Ross K Maddox;Zhiyao Duan;Chenliang Xu
Lele Chen;Haitian Zheng;Ross K Maddox;Zhiyao Duan;Chenliang Xu
中科院分区:
其他
文献类型:
--
作者:
Lele Chen;Haitian Zheng;Ross K Maddox;Zhiyao Duan;Chenliang Xu

文献摘要

相似文献

对以另一种模态为条件的移动人脸/身体的动态进行建模是计算机视觉中的一个基本问题,其中应用范围从音频到视频生成[3]到文本到视频生成以及图像到图像/视频生成[7]。本文考虑这样的任务:给定目标面部图像和任意语音音频记录,生成目标主体的照片般逼真的说话面部,该说话面部具有自然的嘴唇同步,同时保持面部图像随时间的平滑过渡(参见图1)。注意,模型应该具有对不同类型的面部(例如,卡通脸、动物脸)和噪声语音条件。解决这一任务对于实现许多应用至关重要,例如,为听力受损的人从电话音频中读取唇语,为电影和游戏生成具有同步面部动作的虚拟角色。静态图像生成和视频生成之间的主要区别是时间依赖性建模。它带来额外挑战的主要原因有两个:人们对任何像素抖动都很敏感(例如,时间不连续性和细微的伪像);它们还对面部运动和语音音频之间的轻微未对准敏感。然而,最近的研究人员[3,2]倾向于将视频生成制定为时间独立的图像生成问题。在本文中,我们提出了一种新的时间GAN结构,它由一个多模态卷积RNN(MMCRNN)的生成器和一个新的基于回归的神经网络结构。通过对时间依赖性进行建模,我们的基于MMCRNN的生成器可以在相邻帧之间产生更平滑的事务。我们的回归为基础的视频结构相结合的序列级(时间)的信息和帧级(像素变化)的信息来评估生成的视频。说话面部生成的另一个挑战是处理各种视觉动态(例如,摄像机角度、头部移动)与语音音频不相关并且因此不能从语音音频推断。这些复杂的动态,如果在像素空间中建模,将导致低质量的视频。例如,在网络视频(例如,LRW和VoxCeleb数据集),说话者在说话时会明显移动。然而,所有最近的照片逼真的说话脸生成方法[3,9]都没有考虑这个问题。在本文中,我们提出了一种分层结构的音频信号告诉哪里改变
Modeling the dynamics of a moving human face/body conditioned on another modality is a fundamental problem in computer vision, where applications are ranging from audio-to-video generation [3] to text-to-video generation and to skeleton-to-image/video generation [7]. This paper considers such a task: given a target face image and an arbitrary speech audio recording, generating a photo-realistic talking face of the target subject saying that speech with natural lip synchronization while maintaining a smooth transition of facial images over time (see Fig. 1). Note that the model should have a robust generalization capability to different types of faces (e.g., cartoon faces, animal faces) and to noisy speech conditions. Solving this task is crucial to enabling many applications, e.g., lip-reading from over-thephone audio for hearing-impaired people, generating virtual characters with synchronized facial movements to speech audio for movies and games. The main difference between still image generation and video generation is temporal-dependency modeling. There are two main reasons why it imposes additional challenges: people are sensitive to any pixel jittering (e.g., temporal discontinuities and subtle artifacts) in a video; they are also sensitive to slight misalignment between facial movements and speech audio. However, recent researchers [3, 2] tended to formulate video generation as a temporally independent image generation problem. In this paper, we propose a novel temporal GAN structure, which consists of a multi-modal convolutional-RNN-based (MMCRNN) generator and a novel regression-based discriminator structure. By modeling temporal dependencies, our MMCRNNbased generator yields smoother transactions between adjacent frames. Our regression-based discriminator structure combines sequence-level (temporal) information and frame-level (pixel variations) information to evaluate the generated video. Another challenge of the talking face generation is to handle various visual dynamics (e.g., camera angles, head movements) that are not relevant to and hence cannot be inferred from speech audio. Those complicated dynamics, if modeled in the pixel space, will result in low-quality videos. For example, in web videos (e.g., LRW and VoxCeleb datasets), speakers move significantly when they are talking. Nonetheless, all the recent photo-realistic talking face generation methods [3, 9] failed to consider this problem. In this paper, we propose a hierarchical structure Audio signal Tell where to change