Practical Cross-modal Manifold Alignment for Robotic Grounded Language Learning

Practical Cross-modal Manifold Alignment for Robotic Grounded Language Learning
复制标题

DOI:
10.1109/cvprw53098.2021.00177
复制
发表时间:
2021-06
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
A. Nguyen;Frank Ferraro;Cynthia Matuszek
A. Nguyen;Frank Ferraro;Cynthia Matuszek
中科院分区:
其他
文献类型:
--
作者:
A. Nguyen;Frank Ferraro;Cynthia Matuszek

文献摘要

相似文献

我们提出了一个跨模态流形对齐程序,利用三重损失,共同学习一致的,多模态嵌入的语言为基础的概念的现实世界的项目。我们的方法通过从RGB深度图像及其自然语言描述中采样锚、正和负数据点的三元组来学习这些嵌入。我们表明,我们的方法可以受益于,但不需要,后处理步骤,如Procrustes分析,在对比我们的一些基线,需要它的合理性能。我们在两个通常用于开发基于机器人的接地语言学习系统的数据集上展示了我们方法的有效性,其中我们的方法在五个评估指标上优于四个基线,包括最先进的方法。
We propose a cross-modality manifold alignment procedure that leverages triplet loss to jointly learn consistent, multi-modal embeddings of language-based concepts of real-world items. Our approach learns these embeddings by sampling triples of anchor, positive, and negative data points from RGB-depth images and their natural language descriptions. We show that our approach can benefit from, but does not require, post-processing steps such as Procrustes analysis, in contrast to some of our baselines which require it for reasonable performance. We demonstrate the effectiveness of our approach on two datasets commonly used to develop robotic-based grounded language learning systems, where our approach outperforms four baselines, including a state-of-the-art approach, across five evaluation metrics.