ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments

ArraMon: A Joint Navigation-Assembly Instruction Interpretation Task in Dynamic Environments
复制标题

DOI:
10.18653/v1/2020.findings-emnlp.348
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Hyounghun Kim;Abhaysinh Zala;Graham Burri;Hao Tan;Mohit Bansal
Hyounghun Kim;Abhaysinh Zala;Graham Burri;Hao Tan;Mohit Bansal
中科院分区:
其他
文献类型:
--
作者:
Hyounghun Kim;Abhaysinh Zala;Graham Burri;Hao Tan;Mohit Bansal

文献摘要

相似文献

对于具身智能体来说,导航是一种重要的能力,但不是一个孤立的目标。agent还需要在到达目标位置后执行特定的任务,例如拾取物体并将它们组装成特定的排列。我们将视觉和语言导航、收集对象的组装和对象引用表达式理解结合起来,创建了一个新的导航和组装联合任务,命名为ARRAMON。在这个任务中,智能体(类似于PokeMON GO玩家)被要求在一个复杂、现实的户外环境中,根据自然语言(英语)指令,逐个寻找和收集不同的目标物体,然后在一个以自我为中心的网格布局环境中,逐个排列收集到的物体。为了支持这项任务,我们实现了一个3D动态环境模拟器,并收集了一个包含人工编写的导航和组装指令以及相应的地面真值轨迹的数据集。我们还通过验证阶段过滤收集到的指令,导致总共有7.7K任务实例(30.8K指令和路径)。我们展示了几个基线模型(集成的和有偏差的)和指标(nDTW、CTC、rPOD和PTC)的结果,模型与人类的巨大表现差距表明我们的任务是具有挑战性的,并为未来的工作提供了广阔的空间。
For embodied agents, navigation is an important ability but not an isolated goal. Agents are also expected to perform specific tasks after reaching the target location, such as picking up objects and assembling them into a particular arrangement. We combine Vision-andLanguage Navigation, assembling of collected objects, and object referring expression comprehension, to create a novel joint navigation-and-assembly task, named ARRAMON. During this task, the agent (similar to a PokeMON GO player) is asked to find and collect different target objects one-by-one by navigating based on natural language (English) instructions in a complex, realistic outdoor environment, but then also ARRAnge the collected objects part-by-part in an egocentric grid-layout environment. To support this task, we implement a 3D dynamic environment simulator and collect a dataset with human-written navigation and assembling instructions, and the corresponding ground truth trajectories. We also filter the collected instructions via a verification stage, leading to a total of 7.7K task instances (30.8K instructions and paths). We present results for several baseline models (integrated and biased) and metrics (nDTW, CTC, rPOD, and PTC), and the large model-human performance gap demonstrates that our task is challenging and presents a wide scope for future work.