A Deep Learning-Based Model for Head and Eye Motion Generation in Three-party Conversations

A Deep Learning-Based Model for Head and Eye Motion Generation in Three-party Conversations
复制标题

基于深度学习的三方对话中头部和眼睛运动生成模型

DOI:
--
复制
发表时间:
2019
期刊:
PACMCGIT
影响因子:
--
通讯作者:
Z. Deng
Z. Deng
中科院分区:
--
文献类型:
--
作者:
Aobo Jin;Qixin Deng;Yuting Zhang;Z. Deng

文献摘要

被引文献

相似文献

在本文中,我们提出了一种基于深度学习的新颖方法,基于新颖的声学语音输入和说话者标记(即每个对话者的说话时间)来生成真实的三方头部和眼睛运动。具体来说,我们首先获取高质量的三方对话运动数据集。然后,根据获取的数据集,我们训练一个基于深度学习的框架,根据语音信号输入自动预测所有对话者眼睛和头部的动态方向。通过结合现有的口型同步和语音驱动的手/身体手势生成算法,我们可以生成逼真的三方对话动画。通过许多实验和比较用户研究,我们证明我们的方法可以根据不同性别和种族的新受试者录制的新颖语音生成逼真的三方头眼运动。
In this paper we propose a novel deep learning based approach to generate realistic three-party head and eye motions based on novel acoustic speech input together with speaker marking (i.e., speaking time for each interlocutor). Specifically, we first acquire a high quality, three-party conversational motion dataset. Then, based on the acquired dataset, we train a deep learning based framework to automatically predict the dynamic directions of both the eyes and heads of all the interlocutors based on speech signal input. Via the combination of existing lip-sync and speech-driven hand/body gesture generation algorithms, we can generate realistic three-party conversational animations. Through many experiments and comparative user studies, we demonstrate that our approach can generate realistic three-party head-and-eye motions based on novel speech recorded on new subjects with different genders and ethnicities.