Learning shared embedding representation of motion and text using contrastive learning
Learning shared embedding representation of motion and text using contrastive learning
复制标题
DOI:
10.1007/s10015-022-00840-0
复制
发表时间:
2022-12
影响因子:
0.9
通讯作者:
Junpei Horie;Wataru Noguchi;H. Iizuka;Masahito Yamamoto
中科院分区:
文献类型:
--
作者:
Junpei Horie;Wataru Noguchi;H. Iizuka;Masahito Yamamoto
Multimodal learning of motion and text tries to find the correspondence between skeletal time-series data acquired by motion capture and the text that describes the motion. In this field, good associations can realize both motion-to-text and text-to-motion applications. However, the previous methods failed to associate motion with text, taking into account details of descriptions, for example, whether to move the left or right arm. In this paper, we propose a motion-text contrastive learning method for making correspondences between motion and text in a shared embedding space. We showed that our model outperforms the previous studies in the task of action recognition. We also qualitatively show that, by using a pre-trained text encoder, our model can perform motion retrieval with detailed correspondences between motion and text.