Speech Driven Talking Face Generation From a Single Image and an Emotion Condition

Speech Driven Talking Face Generation From a Single Image and an Emotion Condition
复制标题

DOI:
10.1109/tmm.2021.3099900
复制
发表时间:
2021-07-26
影响因子:
7.3
通讯作者:
Duan, Zhiyao
Duan, Zhiyao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Eskimez, Sefik Emre;Zhang, You;Duan, Zhiyao

文献摘要

被引文献

相似文献

视觉情感表达在视听言语交际中起着重要的作用。在这项工作中,我们提出了一种新的方法来渲染语音驱动的说话脸生成中的视觉情感表达。具体来说,我们设计了一个端到端的说话脸生成系统,该系统以语音、单个面部图像和分类情感标签为输入,生成与语音同步并表达条件情感的说话脸视频。对图像质量、视听同步和视觉情感表达的客观评价表明,所提出的系统优于最先进的基线系统。视觉情感表达和视频真实感的主观评价也证明了该系统的优越性。此外,我们使用音频和视觉模式中不匹配的情感生成视频进行人类情感识别试点研究。结果表明,在该任务中,人类对视觉模态的反应比听觉模态更显著。
Visual emotion expression plays an important role in audiovisual speech communication. In this work, we propose a novel approach to rendering visual emotion expression in speech-driven talking face generation. Specifically, we design an end-to-end talking face generation system that takes a speech utterance, a single face image, and a categorical emotion label as input to render a talking face video synchronized with the speech and expressing the conditioned emotion. Objective evaluation on image quality, audiovisual synchronization, and visual emotion expression shows that the proposed system outperforms a state-of-the-art baseline system. Subjective evaluation of visual emotion expression and video realness also demonstrates the superiority of the proposed system. Furthermore, we conduct a human emotion recognition pilot study using generated videos with mismatched emotions among the audio and visual modalities. Results show that humans respond to the visual modality more significantly than the audio modality on this task.