The IVI Lab entry to the GENEA Challenge 2022 – A Tacotron2 Based Method for Co-Speech Gesture Generation With Locality-Constraint Attention Mechanism

The IVI Lab entry to the GENEA Challenge 2022 – A Tacotron2 Based Method for Co-Speech Gesture Generation With Locality-Constraint Attention Mechanism
复制标题

IVI 实验室参加 2022 年 GENEA 挑战赛 — 基于 Tacotron2 的具有局部约束注意机制的共同语音手势生成方法

DOI:
10.1145/3536221.3558060
复制
发表时间:
2022
期刊:
GENEA Challenge 2022
影响因子:
--
通讯作者:
Kapadia, Mubbasir
Kapadia, Mubbasir
中科院分区:
--
文献类型:
--
作者:
Chang, Che-Jui;Zhang, Sen;Kapadia, Mubbasir

文献摘要

参考文献

被引文献

相似文献

本文介绍了IVI实验室参加2022年GENEA挑战的情况。我们将手势生成问题表述为一个序列到序列的转换任务,其中文本、音频和说话者身份作为输入,身体运动作为输出。我们使用Tacotron2架构作为我们的主干,并使用位置约束注意机制来引导解码器从邻近的潜在特征中学习依赖关系。GENEA Challenge 2022发布的集体评估表明,我们的两个全身和上体赛道(FSH和USK)在这两个主观指标上的统计表现都优于音频驱动和文本驱动基线。值得注意的是,我们的全身作品在所有提交的作品中获得了最高的语言得体性(60.5%匹配)。我们还进行了客观的评估,比较我们的运动加速度和抽搐与两个自回归基线。结果表明,我们生成的手势的运动分布更接近于自然手势的分布。
This paper describes the IVI Lab entry to the GENEA Challenge 2022. We formulate the gesture generation problem as a sequence-to-sequence conversion task with text, audio, and speaker identity as inputs and the body motion as the output. We use the Tacotron2 architecture as our backbone with the locality-constraint attention mechanism that guides the decoder to learn the dependencies from the neighboring latent features. The collective evaluation released by GENEA Challenge 2022 indicates that our two entries (FSH and USK) for the full body and upper body tracks statistically outperform the audio-driven and text-driven baselines on both two subjective metrics. Remarkably, our full-body entry receives the highest speech appropriateness (60.5% matched) among all submitted entries. We also conduct an objective evaluation to compare our motion acceleration and jerk with two autoregressive baselines. The result indicates that the motion distribution of our generated gestures is much closer to the distribution of natural gestures.
机器人学习社交技能:人形机器人协同语音手势生成的端到端学习
DOI: --
发表时间: 2018
期刊: IEEE International Conference on Robotics and Automation
影响因子: --
作者:
Youngwoo Yoon;Woo;Minsu Jang;Jaeyeon Lee;Jaehong Kim;Geehyuk Lee
通讯作者: Geehyuk Lee
DOI: 10.1109/iccv48922.2021.01110
发表时间: 2021-08
期刊: 2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子: --
作者:
Jing Li;Di Kang;Wenjie Pei;Xuefei Zhe;Ying Zhang;Zhenyu He-;Linchao Bao
通讯作者: Jing Li;Di Kang;Wenjie Pei;Xuefei Zhe;Ying Zhang;Zhenyu He-;Linchao Bao
从视频中学习语音驱动的 3D 对话手势
DOI: --
发表时间: 2021
期刊: International Conference on Intelligent Virtual Agents
影响因子: --
作者:
I. Habibie;Weipeng Xu;Dushyant Mehta;Lingjie Liu;H. Seidel;Gerard Pons;Mohamed A. Elgharib;C. Theobalt
通讯作者: C. Theobalt
从文本生成连贯的自发语音和手势
DOI: --
发表时间: 2020
期刊: International Conference on Intelligent Virtual Agents
影响因子: --
作者:
Simon Alexanderson;Éva Székely;G. Henter;Taras Kucherenko;J. Beskow
通讯作者: J. Beskow
DOI: 10.1109/thms.2022.3149173
发表时间: 2021
影响因子: 3.6
作者:
Pieter Wolfert;Nicole L. Robinson;Tony Belpaeme
通讯作者: Tony Belpaeme