PRIMAL: Pathfinding via Reinforcement and Imitation Multi-Agent Learning

PRIMAL: Pathfinding via Reinforcement and Imitation Multi-Agent Learning
复制标题

DOI:
10.1109/lra.2019.2903261
复制
发表时间:
2019-07-01
影响因子:
5.2
通讯作者:
Choset, Howie
Choset, Howie
中科院分区:
计算机科学2区
文献类型:
--
作者:
Sartoretti, Guillaume;Kerr, Justin;Choset, Howie

文献摘要

被引文献

相似文献

多智能体路径查找(MAPF)是许多大规模、真实世界机器人部署的重要组成部分,从空中集群到仓库自动化。然而,尽管社区不断努力,大多数最先进的MAPF规划者仍然依赖于集中规划和规模不足,超过几百个代理。这种规划方法不适合现实世界的部署,其中噪音和不确定性通常需要在线重新计算路径,而当规划时间以秒到分钟为单位时,这是不可能的。我们提出了PRIMAL,一个新的框架MAPF,结合强化和模仿学习,教完全分散的政策,其中代理人反应性地规划路径在线部分可观察的世界,同时表现出隐式协调。该框架扩展了我们以前的工作,通过在训练过程中引入专家MAPF规划器的演示,以及仔细的奖励塑造和环境采样,来进行协作策略的分布式学习。一旦学习,所产生的策略可以复制到任何数量的代理上,并自然地扩展到不同的团队规模和世界维度。我们目前的结果随机世界多达1024代理和比较成功率对国家的最先进的MAPF规划。最后,我们通过实验验证了学到的政策,在一个混合仿真的工厂样机,涉及真实的世界和模拟机器人。
Multi-agent path finding (MAPF) is an essential component of many large-scale, real-world robot deployments, from aerial swarms to warehouse automation. However, despite the community's continued efforts, most state-of-the-art MAPF planners still rely on centralized planning and scale poorly past a few hundred agents. Such planning approaches are maladapted to real-world deployments, where noise and uncertainty often require paths be recomputed online, which is impossible when planning times are in seconds to minutes. We present PRIMAL, a novel framework for MAPF that combines reinforcement and imitation learning to teach fully decentralized policies, where agents reactively plan paths online in a partially observable world while exhibiting implicit coordination. This framework extends our previous work on distributed learning of collaborative policies by introducing demonstrations of an expert MAPF planner during training, as well as careful reward shaping and environment sampling. Once learned, the resulting policy can be copied onto any number of agents and naturally scales to different team sizes and world dimensions. We present results on randomized worlds with up to 1024 agents and compare success rates against state-of-the-art MAPF planners. Finally, we experimentally validate the learned policies in a hybrid simulation of a factory mockup, involving both real world and simulated robots.