Learning Simultaneous Navigation and Construction in Grid Worlds

Learning Simultaneous Navigation and Construction in Grid Worlds
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
ArXiv
影响因子:
--
通讯作者:
Wenyu Han;Haoran Wu;Eisuke Hirota;Alexander Gao;Lerrel Pinto;L. Righetti;Chen Feng
Wenyu Han;Haoran Wu;Eisuke Hirota;Alexander Gao;Lerrel Pinto;L. Righetti;Chen Feng
中科院分区:
其他
文献类型:
--
作者:
Wenyu Han;Haoran Wu;Eisuke Hirota;Alexander Gao;Lerrel Pinto;L. Righetti;Chen Feng

文献摘要

相似文献

我们提议研究一种新的学习任务——移动构建,使智能体能够在1/2/3维网格世界中构建设计好的结构,同时在相同的不断变化的环境中导航。与现有的机器人学习任务(如视觉导航和物体操作)不同,由于精确定位和战略构建规划之间的相互依存关系,这项任务具有挑战性。为了基于深度强化学习(RL)寻求针对这个部分可观测马尔可夫决策过程(POMDP)的通用且自适应的解决方案,我们在这个动态网格世界中设计了一个具有显式循环位置估计的深度循环Q网络(DRQN)。我们大量的实验表明,在Q学习之前对这个位置估计模块进行预训练,可以显著提高由交并比分数衡量的构建性能,在我们对各种基准(包括无模型和基于模型的强化学习、一个手工制作的基于同时定位与地图构建(SLAM)的策略以及人类玩家)的测试中取得了最佳结果。我们的代码可在以下网址获取:https://ai4ce.github.io/SNAC/
We propose to study a new learning task, mobile construction , to enable an agent to build designed structures in 1/2/3D grid worlds while navigating in the same evolving environments. Unlike existing robot learning tasks such as visual navigation and object manipulation, this task is challenging because of the interdependence between accurate localization and strategic construction planning. In pursuit of generic and adaptive solutions to this partially observable Markov decision process (POMDP) based on deep reinforcement learning (RL), we design a Deep Recurrent Q-Network (DRQN) with explicit recurrent position estimation in this dynamic grid world. Our extensive experiments show that pre-training this position estimation module before Q-learning can significantly improve the construction performance measured by the intersection-over-union score, achieving the best results in our benchmark of various baselines including model-free and model-based RL, a handcrafted SLAM-based policy, and human players. Our code is available at: https://ai4ce.github.io/SNAC/ .