Evaluation of text-to-gesture generation model using convolutional neural network

Evaluation of text-to-gesture generation model using convolutional neural network
复制标题

使用卷积神经网络评估文本到手势生成模型

DOI:
10.1016/j.neunet.2022.03.041
复制
发表时间:
2022
期刊:
影响因子:
7.8
通讯作者:
Shinichi Shirakawa
Shinichi Shirakawa
中科院分区:
计算机科学1区
文献类型:
--
作者:
Eiichi Asakawa;Naoshi Kaneko;Dai Hasegawa;Shinichi Shirakawa

文献摘要

相似文献

对话手势在实现与虚拟代理和机器人的自然交互方面起着至关重要的作用。数据驱动的方法,如深度学习和机器学习,在构建手势生成模型方面很有前途,该模型自动为语音或口语文本提供手势运动。本文利用卷积神经网络对基于深度学习的语音手势生成模型进行了实验分析。该模型以一系列口语单词为输入,输出代表对话手势运动的2D关节坐标序列。我们通过在现有的数据集上添加文本信息来准备一个由手势、动作和口语文本组成的数据集,并使用特定说话人的数据来训练模型。通过用户感知研究,将生成的手势的质量与现有的语音到手势生成模型的质量进行比较。主观评价表明,该模型的性能与现有的语音到手势生成模型相当或更好。此外,我们还研究了数据清洗和损失函数选择在文本到手势生成模型中的重要性。我们进一步考察了模型在说话人之间的可转移性。实验结果表明,该模型具有较好的模型可移植性。最后,我们展示了文本到手势生成模型即使在使用转换器架构时也可以生成高质量的手势。
Conversational gestures have a crucial role in realizing natural interactions with virtual agents and robots. Data-driven approaches, such as deep learning and machine learning, are promising in constructing the gesture generation model, which automatically provides the gesture motion for speech or spoken texts. This study experimentally analyzes a deep learning-based gesture generation model from spoken text using a convolutional neural network. The proposed model takes a sequence of spoken words as the input and outputs a sequence of 2D joint coordinates representing the conversational gesture motion. We prepare a dataset consisting of gesture motions and spoken texts by adding text information to an existing dataset and train the models using specific speaker’s data. The quality of the generated gestures is compared with those from an existing speech-to-gesture generation model through a user perceptual study. The subjective evaluation shows that the model performance is comparable or superior to those by the existing speech-to-gesture generation model. In addition, we investigate the importance of data cleansing and loss function selection in the text-to-gesture generation model. We further examine the model transferability between speakers. The experimental results demonstrate successful model transferability of the proposed model. Finally, we show that the text-to-gesture generation model can produce good quality gestures even when using a transformer architecture.