ASSISTER: Assistive Navigation via Conditional Instruction Generation

ASSISTER: Assistive Navigation via Conditional Instruction Generation
复制标题

DOI:
10.1007/978-3-031-20059-5_16
复制
发表时间:
2022
影响因子:
4.7
通讯作者:
Zanming Huang;Zhongkai Shangguan;Jimuyang Zhang;Gilad Bar;M. Boyd;Eshed Ohn-Bar
Zanming Huang;Zhongkai Shangguan;Jimuyang Zhang;Gilad Bar;M. Boyd;Eshed Ohn-Bar
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zanming Huang;Zhongkai Shangguan;Jimuyang Zhang;Gilad Bar;M. Boyd;Eshed Ohn-Bar

文献摘要

相似文献

我们介绍了一种新的视觉和语言导航(VLN)的学习任务,以提供实时指导的盲人追随者位于复杂的动态导航方案。为了探索实时信息的需求和基本挑战,在我们的新的建模任务,我们首先收集一个多模态的现实世界的基准与原位定位和移动性(O &M)的指导。随后,我们利用现实世界的研究,通知设计一个更大规模的仿真基准,从而使全面分析当前VLN模型的局限性。出于视觉O &M指南如何无缝和安全地支持视觉障碍者在导航任务合作时的意识,我们提出了ASSISTER,一个可以体现这种有效指导的模仿学习代理。所提出的辅助VLN代理的条件是导航目标和命令,用于生成与周围的视觉场景相一致的教学句子,同时也仔细考虑立即辅助导航任务。总而言之,我们引入的评估和培训框架朝着下一代无缝,类人辅助代理的可扩展发展迈出了一步。
We introduce a novel vision-and-language navigation (VLN) task of learning to provide real-time guidance to a blind follower situated in complex dynamic navigation scenarios. Towards exploring real-time information needs and fundamental challenges in our novel modeling task, we first collect a multi-modal real-world benchmark with in-situ Orientation and Mobility (O &M) instructional guidance. Subsequently, we leverage the real-world study to inform the design of a larger-scale simulation benchmark, thus enabling comprehensive analysis of limitations in current VLN models. Motivated by how sighted O &M guides seamlessly and safely support the awareness of individuals with visual impairments when collaborating on navigation tasks, we present ASSISTER, an imitation-learned agent that can embody such effective guidance. The proposed assistive VLN agent is conditioned on navigational goals and commands for generating instructional sentences that are coherent with the surrounding visual scene, while also carefully accounting for the immediate assistive navigation task. Altogether, our introduced evaluation and training framework takes a step towards scalable development of the next generation of seamless, human-like assistive agents.