FollowNet: Robot Navigation by Following Natural Language Directions with Deep Reinforcement Learning

FollowNet: Robot Navigation by Following Natural Language Directions with Deep Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2018-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Pararth Shah;Marek Fiser;Aleksandra Faust;J. Kew;Dilek Z. Hakkani-Tür
Pararth Shah;Marek Fiser;Aleksandra Faust;J. Kew;Dilek Z. Hakkani-Tür
中科院分区:
其他
文献类型:
--
作者:
Pararth Shah;Marek Fiser;Aleksandra Faust;J. Kew;Dilek Z. Hakkani-Tür

文献摘要

被引文献

相似文献

理解和遵循人类提供的方向可以使机器人在未知情况下有效导航。我们提出了FollowNet,一种端到端可区分的神经结构,用于学习多模式导航策略。FollowNet将自然语言指令以及视觉和深度输入映射到运动原语。FollowNet使用以其视觉和深度输入为条件的注意力机制来处理指令,以便在执行导航任务时专注于命令的相关部分。深度强化学习(RL)是一种稀疏奖励,它同时学习状态表征、注意函数和控制策略。我们在复杂的自然语言方向的数据集上评估我们的代理,这些方向引导代理通过丰富和现实的模拟房屋数据集。我们证明,FollowNet代理学习执行以前未见过的用类似词汇描述的指令,并成功地沿着训练过程中未遇到的路径导航。与没有注意机制的基线模型相比,该代理显示出30%的改进,在新指令下的成功率为52%。
Understanding and following directions provided by humans can enable robots to navigate effectively in unknown situations. We present FollowNet, an end-to-end differentiable neural architecture for learning multi-modal navigation policies. FollowNet maps natural language instructions as well as visual and depth inputs to locomotion primitives. FollowNet processes instructions using an attention mechanism conditioned on its visual and depth input to focus on the relevant parts of the command while performing the navigation task. Deep reinforcement learning (RL) a sparse reward learns simultaneously the state representation, the attention function, and control policies. We evaluate our agent on a dataset of complex natural language directions that guide the agent through a rich and realistic dataset of simulated homes. We show that the FollowNet agent learns to execute previously unseen instructions described with a similar vocabulary, and successfully navigates along paths not encountered during training. The agent shows 30% improvement over a baseline model without the attention mechanism, with 52% success rate at novel instructions.