Generating coherent spontaneous speech and gesture from text

Generating coherent spontaneous speech and gesture from text
复制标题

从文本生成连贯的自发语音和手势

DOI:
--
复制
发表时间:
2020
期刊:
International Conference on Intelligent Virtual Agents
影响因子:
--
通讯作者:
J. Beskow
J. Beskow
中科院分区:
--
文献类型:
--
作者:
Simon Alexanderson;Éva Székely;G. Henter;Taras Kucherenko;J. Beskow

文献摘要

被引文献

相似文献

具身的人类交流包括语言(言语)和非语言信息(例如,手势和头部运动)。机器学习的最新进展大大改进了生成这两种类型数据的合成版本的技术:在语音方面,文本到语音系统现在能够使用无脚本语音音频作为源材料生成高度令人信服的、听起来很自然的语音。在运动方面,概率运动生成方法现在可以合成生动逼真的语音驱动的3D手势。在本文中,我们首次将这两种最先进的技术以一种连贯的方式结合在一起。具体来说,我们展示了一个概念验证系统,该系统在单扬声器音频和动作捕捉数据集上训练,能够从文本输入中同时生成语音和全身手势。与之前的联合语音和手势生成方法相比,我们通过对来自同一个人的自发语音记录进行训练的语音合成来生成全身手势作为动作捕获数据。我们通过可视化手势空间和文本-语音-手势对齐以及演示视频来说明我们的结果。
Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions of both of these types of data: On the speech side, text-to-speech systems are now able to generate highly convincing, spontaneous-sounding speech using unscripted speech audio as the source material. On the motion side, probabilistic motion-generation methods can now synthesise vivid and lifelike speech-driven 3D gesticulation. In this paper, we put these two state-of-the-art technologies together in a coherent fashion for the first time. Concretely, we demonstrate a proof-of-concept system trained on a single-speaker audio and motion-capture dataset, that is able to generate both speech and full-body gestures together from text input. In contrast to previous approaches for joint speech-and-gesture generation, we generate full-body gestures from speech synthesis trained on recordings of spontaneous speech from the same person as the motion-capture data. We illustrate our results by visualising gesture spaces and textspeech-gesture alignments, and through a demonstration video.