Speaker-Follower Models for Vision-and-Language Navigation

Speaker-Follower Models for Vision-and-Language Navigation
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Daniel Fried;Ronghang Hu;Volkan Cirik;Anna Rohrbach;Jacob Andreas;Louis-Philippe Morency;Taylor Berg-Kirkpatrick;Kate Saenko;D. Klein;Trevor Darrell
Daniel Fried;Ronghang Hu;Volkan Cirik;Anna Rohrbach;Jacob Andreas;Louis-Philippe Morency;Taylor Berg-Kirkpatrick;Kate Saenko;D. Klein;Trevor Darrell
中科院分区:
其他
文献类型:
--
作者:
Daniel Fried;Ronghang Hu;Volkan Cirik;Anna Rohrbach;Jacob Andreas;Louis-Philippe Morency;Taylor Berg-Kirkpatrick;Kate Saenko;D. Klein;Trevor Darrell

文献摘要

被引文献

相似文献

自然语言指令引导的导航对指令跟随者提出了一个具有挑战性的推理问题。自然语言指令通常只识别一些高级决策和地标,而不是完整的低级运动行为;大部分缺失的信息必须基于感知上下文来推断。在机器学习环境中,这是双重挑战:很难收集足够的注释数据来从头开始学习这个推理过程,并且也很难使用通用序列模型来实现推理过程。在这里,我们描述了一种视觉和语言导航的方法,该方法通过嵌入式扬声器模型解决了这两个问题。我们使用这个说话人模型来(1)合成新的指令以进行数据增强,以及(2)实现语用推理,评估候选动作序列解释指令的程度。这两个步骤都得到了全景动作空间的支持,该动作空间反映了人类生成的指令的粒度。实验表明,这种方法的所有三个组成部分-扬声器驱动的数据增强,务实的推理和全景动作空间-显着提高了基线指令跟随器的性能,在标准基准测试中,成功率比现有的最佳方法提高了一倍以上。
Navigation guided by natural language instructions presents a challenging reasoning problem for instruction followers. Natural language instructions typically identify only a few high-level decisions and landmarks rather than complete low-level motor behaviors; much of the missing information must be inferred based on perceptual context. In machine learning settings, this is doubly challenging: it is difficult to collect enough annotated data to enable learning of this reasoning process from scratch, and also difficult to implement the reasoning process using generic sequence models. Here we describe an approach to vision-and-language navigation that addresses both these issues with an embedded speaker model. We use this speaker model to (1) synthesize new instructions for data augmentation and to (2) implement pragmatic reasoning, which evaluates how well candidate action sequences explain an instruction. Both steps are supported by a panoramic action space that reflects the granularity of human-generated instructions. Experiments show that all three components of this approach---speaker-driven data augmentation, pragmatic reasoning and panoramic action space---dramatically improve the performance of a baseline instruction follower, more than doubling the success rate over the best existing approach on a standard benchmark.