Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders

Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders
复制标题

DOI:
10.1109/iccv48922.2021.01110
复制
发表时间:
2021-08
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Jing Li;Di Kang;Wenjie Pei;Xuefei Zhe;Ying Zhang;Zhenyu He-;Linchao Bao
Jing Li;Di Kang;Wenjie Pei;Xuefei Zhe;Ying Zhang;Zhenyu He-;Linchao Bao
中科院分区:
其他
文献类型:
--
作者:
Jing Li;Di Kang;Wenjie Pei;Xuefei Zhe;Ying Zhang;Zhenyu He-;Linchao Bao

文献摘要

被引文献

相似文献

由于音频和身体运动之间固有的一对多映射,从语音音频生成会话手势是具有挑战性的。传统的CNN/RNN假设一对一映射,因此倾向于预测所有可能的目标运动的平均值,从而导致推理过程中出现平淡/无聊的运动。为了克服这个问题,我们提出了一种新的条件变分自动编码器(VAE),明确的模型一对多音频到运动的映射,通过分裂的跨模态的潜在代码到共享代码和特定于运动的代码。共享代码主要模拟音频和运动之间的强相关性(例如同步的音频和运动节拍),而运动特定代码捕获独立于音频的各种运动信息。然而,将潜在代码分成两部分给VAE模型带来了训练困难。一个映射网络,促进随机采样沿着与其他技术,包括放松运动损失,自行车约束,和多样性损失的设计,以更好地训练VAE。在3D和2D运动数据集上的实验验证了我们的方法比最先进的方法在定量和定性上产生更真实和多样化的运动。最后,我们证明了我们的方法可以很容易地用于生成运动序列与用户指定的运动剪辑上的时间轴。代码和更多结果在https://jingli513.github.io/audio2gestures。
Generating conversational gestures from speech audio is challenging due to the inherent one-to-many mapping be-tween audio and body motions. Conventional CNNs/RNNs assume one-to-one mapping, and thus tend to predict the average of all possible target motions, resulting in plain/boring motions during inference. In order to over-come this problem, we propose a novel conditional variational autoencoder (VAE) that explicitly models one-to-many audio-to-motion mapping by splitting the cross-modal latent code into shared code and motion-specific code. The shared code mainly models the strong correlation between audio and motion (such as the synchronized audio and motion beats), while the motion-specific code captures diverse motion information independent of the audio. However, splitting the latent code into two parts poses training difficulties for the VAE model. A mapping network facilitating random sampling along with other techniques including relaxed motion loss, bicycle constraint, and diversity loss are designed to better train the VAE. Experiments on both 3D and 2D motion datasets verify that our method generates more realistic and diverse motions than state-of-the-art methods, quantitatively and qualitatively. Finally, we demonstrate that our method can be readily used to generate motion sequences with user-specified motion clips on the timeline. Code and more results are at https://jingli513.github.io/audio2gestures.