A deep learning approach for generalized speech animation

A deep learning approach for generalized speech animation
复制标题

DOI:
10.1145/3072959.3073699
复制
发表时间:
2017-07
期刊:
ACM Transactions on Graphics (TOG)
影响因子:
--
通讯作者:
Sarah L. Taylor;Taehwan Kim;Yisong Yue;Moshe Mahler;James Krahe;Anastasio Garcia Rodriguez;J. Hodgins
Sarah L. Taylor;Taehwan Kim;Yisong Yue;Moshe Mahler;James Krahe;Anastasio Garcia Rodriguez;J. Hodgins
中科院分区:
其他
文献类型:
--
作者:
Sarah L. Taylor;Taehwan Kim;Yisong Yue;Moshe Mahler;James Krahe;Anastasio Garcia Rodriguez;J. Hodgins

文献摘要

被引文献

相似文献

我们介绍了一种简单有效的深度学习方法,可以自动生成自然的语音动画,以支持输入语音。我们的方法使用滑动窗口预测器,学习任意的非线性映射从音素标签输入序列的嘴部运动的方式,准确地捕捉自然运动和视觉协同发音效果。我们的深度学习方法具有几个有吸引力的特性:它实时运行,需要最少的参数调整,很好地推广到新的输入语音序列,很容易编辑以创建风格化和情感化的语音,并且与现有的动画重定向方法兼容。我们工作的一个重点是开发一种有效的语音动画方法,可以很容易地集成到现有的生产流水线。我们提供了端到端方法的详细描述,包括机器学习设计决策。广义语音动画的结果,证明了在各种角色和声音,包括唱歌和外语输入的动画剪辑范围广泛。我们的方法也可以产生按需语音动画实时从用户的语音输入。
We introduce a simple and effective deep learning approach to automatically generate natural looking speech animation that synchronizes to input speech. Our approach uses a sliding window predictor that learns arbitrary nonlinear mappings from phoneme label input sequences to mouth movements in a way that accurately captures natural motion and visual coarticulation effects. Our deep learning approach enjoys several attractive properties: it runs in real-time, requires minimal parameter tuning, generalizes well to novel input speech sequences, is easily edited to create stylized and emotional speech, and is compatible with existing animation retargeting approaches. One important focus of our work is to develop an effective approach for speech animation that can be easily integrated into existing production pipelines. We provide a detailed description of our end-to-end approach, including machine learning design decisions. Generalized speech animation results are demonstrated over a wide range of animation clips on a variety of characters and voices, including singing and foreign language input. Our approach can also generate on-demand speech animation in real-time from user speech input.