Learning Bidirectional Translation Between Descriptions and Actions With Small Paired Data

Learning Bidirectional Translation Between Descriptions and Actions With Small Paired Data
复制标题

DOI:
10.1109/lra.2022.3196159
复制
发表时间:
2022-03
影响因子:
5.2
通讯作者:
M. Toyoda;Kanata Suzuki;Yoshihiko Hayashi;T. Ogata
M. Toyoda;Kanata Suzuki;Yoshihiko Hayashi;T. Ogata
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. Toyoda;Kanata Suzuki;Yoshihiko Hayashi;T. Ogata

文献摘要

相似文献

本研究使用来自不同模态的小配对数据实现了描述和动作之间的双向翻译。相互生成描述和动作的能力对于机器人在日常生活中与人类合作至关重要,这通常需要一个大型数据集来维护两种模态数据的全面配对。然而,一个配对的数据集是昂贵的建设和难以收集。为了解决这个问题,本研究提出了一个两阶段的双向翻译训练方法。在所提出的方法中,我们训练循环自编码器(RAE)用于具有大量非配对数据的描述和动作。然后,我们对整个模型进行微调,以使用小的配对数据绑定它们的中间表示。由于用于预训练的数据不需要配对,因此可以使用仅行为数据或大型语言语料库。我们使用由运动捕获动作和描述组成的配对数据集对我们的方法进行了实验评估。结果表明,即使要训练的配对数据量很小,我们的方法也表现良好。每个RAE的中间表示的可视化显示,类似的动作被编码在一个聚类的位置和相应的特征向量很好地对齐。
This study achieved bidirectional translation between descriptions and actions using small paired data from different modalities. The ability to mutually generate descriptions and actions is essential for robots to collaborate with humans in their daily lives, which generally requires a large dataset that maintains comprehensive pairs of both modality data. However, a paired dataset is expensive to construct and difficult to collect. To address this issue, this study proposes a two-stage training method for bidirectional translation. In the proposed method, we train recurrent autoencoders (RAEs) for descriptions and actions with a large amount of non-paired data. Then, we fine-tune the entire model to bind their intermediate representations using small paired data. Because the data used for pre-training do not require pairing, behavior-only data or a large language corpus can be used. We experimentally evaluated our method using a paired dataset consisting of motion-captured actions and descriptions. The results showed that our method performed well, even when the amount of paired data to train was small. The visualization of the intermediate representations of each RAE showed that similar actions were encoded in a clustered position and the corresponding feature vectors were well aligned.