Cross-modal Map Learning for Vision and Language Navigation

Cross-modal Map Learning for Vision and Language Navigation
复制标题

DOI:
10.1109/cvpr52688.2022.01502
复制
发表时间:
2022-03
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
G. Georgakis;Karl Schmeckpeper;Karan Wanchoo;Soham Dan;E. Miltsakaki;D. Roth;Kostas Daniilidis
G. Georgakis;Karl Schmeckpeper;Karan Wanchoo;Soham Dan;E. Miltsakaki;D. Roth;Kostas Daniilidis
中科院分区:
其他
文献类型:
--
作者:
G. Georgakis;Karl Schmeckpeper;Karan Wanchoo;Soham Dan;E. Miltsakaki;D. Roth;Kostas Daniilidis

文献摘要

相似文献

我们考虑视觉和语言导航(VLN)的问题。当前大多数 VLN 方法都是使用非结构化记忆(例如 LSTM)或对智能体的自我中心观察使用跨模式注意力进行端到端训练。与其他作品相比,我们的主要见解是,当语言和视觉之间的关联出现在明确的空间表征中时,这种关联会更强。在这项工作中,我们提出了一种用于视觉和语言导航的跨模式地图学习模型,该模型首先学习在以自我为中心的地图上预测观察和未观察区域的自上而下语义,然后将通往目标的路径预测为一组路径点。在这两种情况下,预测都是通过跨模式注意机制由语言通知的。我们通过实验测试了语言驱动导航可以在给定地图的情况下解决的基本假设,然后在完整的 VLN-CE 基准测试中展示具有竞争力的结果。
We consider the problem of Vision-and-Language Navigation (VLN). The majority of current methods for VLN are trained end-to-end using either unstructured memory such as LSTM, or using cross-modal attention over the egocentric observations of the agent. In contrast to other works, our key insight is that the association between language and vision is stronger when it occurs in explicit spatial representations. In this work, we propose a cross-modal map learning model for vision-and-language navigation that first learns to predict the top-down semantics on an egocentric map for both observed and unobserved regions, and then predicts a path towards the goal as a set of way-points. In both cases, the prediction is informed by the language through cross-modal attention mechanisms. We experimentally test the basic hypothesis that language-driven navigation can be solved given a map, and then show competitive results on the full VLN-CE benchmark.