Learning shared embedding representation of motion and text using contrastive learning

Learning shared embedding representation of motion and text using contrastive learning
复制标题

DOI:
10.1007/s10015-022-00840-0
复制
发表时间:
2022-12
影响因子:
0.9
通讯作者:
Junpei Horie;Wataru Noguchi;H. Iizuka;Masahito Yamamoto
Junpei Horie;Wataru Noguchi;H. Iizuka;Masahito Yamamoto
中科院分区:
--
文献类型:
--
作者:
Junpei Horie;Wataru Noguchi;H. Iizuka;Masahito Yamamoto

文献摘要

相似文献

运动和文本的多模式学习试图找到通过运动捕获获取的骨骼时间序列数据和描述运动的文本之间的对应关系。在这一领域,良好的关联可以实现运动到文本和文本到运动的应用。然而,以前的方法未能将运动与文本相关联,考虑到描述的细节,例如,是否移动左臂或右臂。在本文中,我们提出了一种运动文本对比学习方法,用于在共享嵌入空间中进行运动和文本的对应。实验结果表明,我们的模型在动作识别任务中的表现优于前人的研究。我们还定性地证明,通过使用预先训练的文本编码器,我们的模型可以执行运动检索,并且运动和文本之间有详细的对应关系。
Multimodal learning of motion and text tries to find the correspondence between skeletal time-series data acquired by motion capture and the text that describes the motion. In this field, good associations can realize both motion-to-text and text-to-motion applications. However, the previous methods failed to associate motion with text, taking into account details of descriptions, for example, whether to move the left or right arm. In this paper, we propose a motion-text contrastive learning method for making correspondences between motion and text in a shared embedding space. We showed that our model outperforms the previous studies in the task of action recognition. We also qualitatively show that, by using a pre-trained text encoder, our model can perform motion retrieval with detailed correspondences between motion and text.