A Modular Vision Language Navigation and Manipulation Framework for Long Horizon Compositional Tasks in Indoor Environment.

A Modular Vision Language Navigation and Manipulation Framework for Long Horizon Compositional Tasks in Indoor Environment.
复制标题

DOI:
10.3389/frobt.2022.930486
复制
发表时间:
2022
影响因子:
3.4
通讯作者:
Sarkar, Soumik
Sarkar, Soumik
中科院分区:
其他
文献类型:
--
作者:
Saha, Homagni;Fotouhi, Fateme;Liu, Qisai;Sarkar, Soumik

文献摘要

被引文献

相似文献

在本文中,我们提出了一个新的框架-MoViLan(模块化视觉和语言)执行视觉接地自然语言指令的日常室内家居任务。虽然已经提出了几个数据驱动的,端到端的学习框架,有针对性的导航任务的基础上的视觉和语言模式,最近的基准数据集的性能揭示了差距,在开发全面的技术,长期的组成任务(涉及操纵和导航)与不同的对象类别,现实的指令和视觉场景与不可逆的状态变化。我们提出了一种模块化方法来处理组合导航和对象交互问题,而不需要严格对齐的视觉和语言训练数据(例如,以专家演示轨迹的形式)。这种方法与该领域传统的端到端技术有很大不同,并且允许使用单独的视觉和语言数据集进行更易于处理的训练过程。具体来说,我们提出了一种新的几何感知映射技术,为杂乱的室内环境,和语言理解模型概括为家庭教学以下。我们证明了一个显着增加的成功率为长期的视野,组成的任务,最近的作品在最近发布的基准数据集-ALFRED。
In this paper we propose a new framework—MoViLan (Modular Vision and Language) for execution of visually grounded natural language instructions for day to day indoor household tasks. While several data-driven, end-to-end learning frameworks have been proposed for targeted navigation tasks based on the vision and language modalities, performance on recent benchmark data sets revealed the gap in developing comprehensive techniques for long horizon, compositional tasks (involving manipulation and navigation) with diverse object categories, realistic instructions and visual scenarios with non reversible state changes. We propose a modular approach to deal with the combined navigation and object interaction problem without the need for strictly aligned vision and language training data (e.g., in the form of expert demonstrated trajectories). Such an approach is a significant departure from the traditional end-to-end techniques in this space and allows for a more tractable training process with separate vision and language data sets. Specifically, we propose a novel geometry-aware mapping technique for cluttered indoor environments, and a language understanding model generalized for household instruction following. We demonstrate a significant increase in success rates for long horizon, compositional tasks over recent works on the recently released benchmark data set -ALFRED.