Noise-Resilient Training Method for Face Landmark Generation From Speech

Noise-Resilient Training Method for Face Landmark Generation From Speech
复制标题

DOI:
10.1109/taslp.2019.2947741
复制
发表时间:
2020-05
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
S. Eskimez;R. Maddox;Chenliang Xu;Z. Duan
S. Eskimez;R. Maddox;Chenliang Xu;Z. Duan
中科院分区:
其他
文献类型:
--
作者:
S. Eskimez;R. Maddox;Chenliang Xu;Z. Duan

文献摘要

相似文献

视觉线索,如嘴唇运动,当可用时,在言语交际中起着重要作用。它们对于听力受损人群或嘈杂环境中的人特别有帮助。当不可用时,具有与输入语音同步地自动生成说话面部的系统将增强语音通信并实现许多新颖的应用。在这篇文章中,我们提出了一个新的系统,可以生成3D说话的脸地标从语音在网上的时尚。我们采用一个神经网络,接受原始波形作为输入。该网络包含具有1D内核的卷积层,并输出面部地标的主动形状模型(ASM)系数。为了促进视频帧之间更平滑的过渡,我们提出了一种具有相同架构的模型变体,但也接受前一帧的ASM系数作为额外的输入。为了科普背景噪声,我们提出了一种新的训练方法,将语音增强的想法在特征级。对地标预测的客观评价表明,该系统在单说话人数据集和多说话人数据集上产生的误差显著小于两种最先进的基线方法。在五种非平稳不可见噪声的语音输入上的实验表明,由于噪声弹性训练方法,系统性能在统计上有显着的改善。最后,主观评估表明,生成的说话的脸有一个显着更令人信服的匹配与输入音频,实现了类似的令人信服的现实主义水平的地面实况地标。
Visual cues such as lip movements, when available, play an important role in speech communication. They are especially helpful for the hearing impaired population or in noisy environments. When not available, having a system to automatically generate talking faces in sync with input speech would enhance speech communication and enable many novel applications. In this article, we present a new system that can generate 3D talking face landmarks from speech in an online fashion. We employ a neural network that accepts the raw waveform as an input. The network contains convolutional layers with 1D kernels and outputs the active shape model (ASM) coefficients of face landmarks. To promote smoother transitions between video frames, we present a variant of the model that has the same architecture but also accepts the previous frame's ASM coefficients as an additional input. To cope with background noise, we propose a new training method to incorporate speech enhancement ideas at the feature level. Objective evaluations on landmark prediction show that the proposed system yields statistically significantly smaller errors than two state-of-the-art baseline methods on both a single-speaker dataset and a multi-speaker dataset. Experiments on noisy speech input with five types of non-stationary unseen noise show statistically significant improvements of the system performance thanks to the noise-resilient training method. Finally, subjective evaluations show that the generated talking faces have a significantly more convincing match with the input audio, achieving a similarly convincing level of realism as the ground-truth landmarks.