Learning Speech-driven 3D Conversational Gestures from Video

Learning Speech-driven 3D Conversational Gestures from Video
复制标题

从视频中学习语音驱动的 3D 对话手势

DOI:
--
复制
发表时间:
2021
期刊:
International Conference on Intelligent Virtual Agents
影响因子:
--
通讯作者:
C. Theobalt
C. Theobalt
中科院分区:
--
文献类型:
--
作者:
I. Habibie;Weipeng Xu;Dushyant Mehta;Lingjie Liu;H. Seidel;Gerard Pons;Mohamed A. Elgharib;C. Theobalt

文献摘要

参考文献

被引文献

相似文献

我们提出了第一种方法,从语音输入合成虚拟角色的同步3D会话身体和手势,以及3D面部和头部动画。我们的算法使用CNN架构,利用面部表情和手势之间的内在相关性。会话肢体手势的合成是一个多模态问题,因为许多相似的手势可能伴随着相同的输入语音。为了在这种情况下合成合理的身体手势,我们训练了一个基于生成对抗网络(GAN)的模型,该模型在与输入音频特征配对时测量生成的3D身体运动序列的合理性。我们还提供了一个新的语料库,其中包含超过33小时的注释数据,这些数据来自于说话的人的野外视频。为此,我们将最先进的单眼方法应用于3D身体和手部姿势估计以及3D面部表现捕获视频语料库。通过这种方式,我们可以训练比以前的算法更多的数量级数据,这些算法诉诸于复杂的工作室内动作捕捉解决方案,从而训练更具表现力的合成算法。我们的实验和用户研究表明,我们的语音合成全3D角色动画的质量是最先进的。
We propose the first approach to synthesize the synchronous 3D conversational body and hand gestures, as well as 3D face and head animations, of a virtual character from speech input. Our algorithm uses a CNN architecture that leverages the inherent correlation between facial expression and hand gestures. Synthesis of conversational body gestures is a multi-modal problem since many similar gestures can plausibly accompany the same input speech. To synthesize plausible body gestures in this setting, we train a Generative Adversarial Network (GAN) based model that measures the plausibility of the generated sequences of 3D body motion when paired with the input audio features. We also contribute a new corpus that contains more than 33 hours of annotated data from in-the-wild videos of talking people. To this end, we apply state-of-the-art monocular approaches for 3D body and hand pose estimation as well as 3D face performance capture to the video corpus. In this way, we can train on orders of magnitude more data than previous algorithms that resort to complex in-studio motion capture solutions, and thereby train more expressive synthesis algorithms. Our experiments and user study show the state-of-the-art quality of our speech-synthesized full 3D character animations.
DOI: 10.1109/tvcg.2018.2868527
发表时间: 2018-09
影响因子: 5.2
作者:
Young-Woon Cha;True Price;Zhen Wei;Xinran Lu;Nicholas Rewkowski;Rohan Chabra;Zihe Qin;Hyounghun Kim;Zhaoqi Su;Yebin Liu;A. Ilie;A. State;Zhenlin Xu;Jan-Michael Frahm;H. Fuchs
通讯作者: Young-Woon Cha;True Price;Zhen Wei;Xinran Lu;Nicholas Rewkowski;Rohan Chabra;Zihe Qin;Hyounghun Kim;Zhaoqi Su;Yebin Liu;A. Ilie;A. State;Zhenlin Xu;Jan-Michael Frahm;H. Fuchs
DOI: 10.1145/3072959.3073699
发表时间: 2017-07
期刊: ACM Transactions on Graphics (TOG)
影响因子: --
作者:
Sarah L. Taylor;Taehwan Kim;Yisong Yue;Moshe Mahler;James Krahe;Anastasio Garcia Rodriguez;J. Hodgins
通讯作者: Sarah L. Taylor;Taehwan Kim;Yisong Yue;Moshe Mahler;James Krahe;Anastasio Garcia Rodriguez;J. Hodgins
由情感合成语音驱动的有意义的头部运动
DOI: 10.1016/j.specom.2017.07.004
发表时间: 2017
影响因子: 3.2
作者:
Sadoughi, Najmeh;Liu, Yang;Busso, Carlos
通讯作者: Busso, Carlos