A Deep Learning-Based Model for Head and Eye Motion Generation in Three-party Conversations
A Deep Learning-Based Model for Head and Eye Motion Generation in Three-party Conversations
复制标题
基于深度学习的三方对话中头部和眼睛运动生成模型
DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Z. Deng
中科院分区:
文献类型:
--
作者:
Aobo Jin;Qixin Deng;Yuting Zhang;Z. Deng
In this paper we propose a novel deep learning based approach to generate realistic three-party head and eye motions based on novel acoustic speech input together with speaker marking (i.e., speaking time for each interlocutor). Specifically, we first acquire a high quality, three-party conversational motion dataset. Then, based on the acquired dataset, we train a deep learning based framework to automatically predict the dynamic directions of both the eyes and heads of all the interlocutors based on speech signal input. Via the combination of existing lip-sync and speech-driven hand/body gesture generation algorithms, we can generate realistic three-party conversational animations. Through many experiments and comparative user studies, we demonstrate that our approach can generate realistic three-party head-and-eye motions based on novel speech recorded on new subjects with different genders and ethnicities.