A Spatio-temporal Learning for Music Conditioned Dance Generation

A Spatio-temporal Learning for Music Conditioned Dance Generation
复制标题

DOI:
10.1145/3536221.3556618
复制
发表时间:
2022-11
期刊:
Proceedings of the 2022 International Conference on Multimodal Interaction
影响因子:
--
通讯作者:
Li Zhou;Yan Luo
Li Zhou;Yan Luo
中科院分区:
其他
文献类型:
--
作者:
Li Zhou;Yan Luo

文献摘要

相似文献

音乐条件下的舞蹈生成,即随音乐起舞,是多模态人体动作合成的一种使用场景。通常情况下,设计符合音乐旋律和节奏的连续动作是一个挑战。本文提出了一种基于位置的编码解码框架,用于运动的时空学习和面向音乐的长期基于骨骼的舞蹈生成。考虑到1分钟视频片段中帧的位置嵌入,首先,我们模块化了一个基于区域注意力的前馈机制来编码音乐特征。其次,基于每个帧的骨架和跨运动帧的联合轨迹,我们形式化了一个图拓扑来表示每个舞蹈序列的时空知识。具体来说,我们提出了一个基于图形卷积网络(GCN)的块来处理运动的长期依赖关系,并利用空间和时间特征。音乐和运动路径都是在位置嵌入方案中完全学习的,并通过重复相应的块来构建。最后,由于舞蹈生成的任务本质上是音乐和动作之间的一致性,我们提出了一种跨模态特征融合的多模态交互和音乐条件下的舞蹈生成。实验结果表明,我们的方法在运动质量和运动-音乐相关度量方面优于最先进的方法。
The music-conditioned dance generation, i.e., dancing to music, is a usage scenario of multi-modality human motion synthesis. Typically, it is a challenge to choreograph continuous motions coinciding with the melody and rhythm of the music. This paper proposes a position-wise encoding-decoding framework for spatio-temporal learning of motions and long-term skeleton-based dance generation oriented on music. Given the positional embedding of the frames in 1-minute video clips, firstly, we modularize a regional attention-based feed-forward mechanism to encode the music features. Secondly, based on the skeleton of each frame and the joint trajectories across motion frames, we formalize a graph topology to represent each dance sequence’s spatial and temporal knowledge. Specifically, we propose a graph convolutional network (GCN) based blocks to process long-term dependencies of motions and leverage the spatial and temporal features. Both music and motion paths are learned fully in positional embedding schemes and constructed by repeating the corresponding blocks. Finally, as the task of dance generation is inherently the consistency between music and motions, we proposed a cross-modality feature fusion for multimodal interaction and music-conditioned dance generation. Experimental results demonstrate that our method outperforms state-of-art methods in motion quality and motion-music correlation metrics.