Expressive Speech-Driven Lip Movements with Multitask Learning

Expressive Speech-Driven Lip Movements with Multitask Learning
复制标题

DOI:
10.1109/fg.2018.00066
复制
发表时间:
2018-05
期刊:
2018 13th IEEE International Conference on Automatic Face & Gesture Recognition (FG 2018)
影响因子:
--
通讯作者:
Najmeh Sadoughi;C. Busso
Najmeh Sadoughi;C. Busso
中科院分区:
其他
文献类型:
--
作者:
Najmeh Sadoughi;C. Busso

文献摘要

相似文献

口面部区域传达一系列信息,包括语音清晰度和情绪。这两个因素增加了对面部运动的限制,创造了不平凡的整合和相互作用。为了为会话主体(CA)产生更具表现力和自然主义的动作,这些因素之间的关系应该被仔细地建模。与基于规则的系统相比,数据驱动模型更适合于这项任务。本文提出了两种深度学习语音驱动结构来整合语音清晰度和情感线索。所提出的方法依赖于多任务学习(MTL)策略,在合成口腔面部动作时,相关的次级任务被共同解决。特别是,我们将情感识别和视位识别作为次要任务。这种方法创建了共享的表征,这些表征产生的行为不仅更接近原始的口腔面部运动,而且比单一任务学习的结果更自然。
The orofacial area conveys a range of information, including speech articulation and emotions. These two factors add constraints to the facial movements, creating non-trivial integrations and interplays. To generate more expressive and naturalistic movements for conversational agents (CAs) the relationship between these factors should be carefully modeled. Data-driven models are more appropriate for this task than rule-based systems. This paper provides two deep learning speech-driven structures to integrate speech articulation and emotional cues. The proposed approaches rely on multitask learning (MTL) strategies, where related secondary tasks are jointly solved when synthesizing orofacial movements. In particular, we evaluate emotion recognition and viseme recognition as secondary tasks. The approach creates shared representations that generate behaviors that not only are closer to the original orofacial movements, but also are perceived more natural than the results from single task learning.