Toward Fine-Grained Talking Face Generation

Toward Fine-Grained Talking Face Generation
复制标题

DOI:
10.1109/tip.2023.3323452
复制
发表时间:
2023-10
影响因子:
10.6
通讯作者:
Zhicheng Sheng;Liqiang Nie;Meng Liu;Yin-wei Wei;Zan Gao
Zhicheng Sheng;Liqiang Nie;Meng Liu;Yin-wei Wei;Zan Gao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhicheng Sheng;Liqiang Nie;Meng Liu;Yin-wei Wei;Zan Gao

文献摘要

相似文献

说话脸生成是在给定参考肖像和音频剪辑时合成嘴唇同步视频的过程。然而,由于以下几个挑战,生成细粒度的谈话视频并不简单:1)捕捉生动的面部表情,例如肌肉运动; 2)确保连续帧之间的平滑过渡;以及3)保留参考肖像的细节。现有的努力只集中在建模刚性嘴唇运动,导致低保真度的视频与抖动的面部肌肉变形。为了解决这些挑战,我们提出了一种新的Fine-grained mOtioN moDel(FROND),由三个组件组成。在第一个部分中,我们采用了两个流编码器来捕获局部面部运动关键点,并将其整体运动上下文作为全局代码嵌入。在第二部分中,我们设计了一个运动估计模块来预测音频驱动的运动。这使得能够在连续轨迹空间中学习局部关键点运动,以实现平滑的时间面部运动。此外,融合局部和全局运动来估计连续密集的运动场,从而产生空间平滑的运动。在第三部分中,我们设计了一个新的隐式图像解码器的隐式神经网络的基础上。该解码器从输入图像中恢复高频信息,从而产生高保真度的说话面部。总之,FROND将面部关键点的运动轨迹细化为连续的密集运动场,随后是充分利用运动的固有平滑度的解码器。我们对基准数据集进行定量和定性模型评估。实验结果表明,我们提出的FROND显着优于几个国家的最先进的基线。
Talking face generation is the process of synthesizing a lip-synchronized video when given a reference portrait and an audio clip. However, generating a fine-grained talking video is nontrivial due to several challenges: 1) capturing vivid facial expressions, such as muscle movements; 2) ensuring smooth transitions between consecutive frames; and 3) preserving the details of the reference portrait. Existing efforts have only focused on modeling rigid lip movements, resulting in low-fidelity videos with jerky facial muscle deformations. To address these challenges, we propose a novel Fine-gRained mOtioN moDel (FROND), consisting of three components. In the first component, we adopt a two-stream encoder to capture local facial movement keypoints and embed their overall motion context as the global code. In the second component, we design a motion estimation module to predict audio-driven movements. This enables the learning of local key point motion in the continuous trajectory space to achieve smooth temporal facial movements. Additionally, the local and global motions are fused to estimate a continuous dense motion field, resulting in spatially smooth movements. In the third component, we devise a novel implicit image decoder based on an implicit neural network. This decoder recovers high-frequency information from the input image, resulting in a high-fidelity talking face. In summary, the FROND refines the motion trajectories of facial keypoints into a continuous dense motion field, which is followed by a decoder that fully exploits the inherent smoothness of the motion. We conduct quantitative and qualitative model evaluations on benchmark datasets. The experimental results show that our proposed FROND significantly outperforms several state-of-the-art baselines.