Multi-Robot Guided Policy Search for Learning Decentralized Swarm Control

Multi-Robot Guided Policy Search for Learning Decentralized Swarm Control
复制标题

DOI:
10.1109/lcsys.2020.3005441
复制
发表时间:
2021-07-01
影响因子:
3
通讯作者:
Guo, Yi
Guo, Yi
中科院分区:
其他
文献类型:
--
作者:
Jiang, Chao;Guo, Yi

文献摘要

被引文献

相似文献

近年来,多机器人学习得到了广泛的研究。开发学习分散控制政策的可证明是正确的算法仍然具有挑战性。在本文中,我们提出了一种基于引导策略搜索的样本高效的多机器人学习方法来学习分散的群体控制策略。该方法采用分布式轨迹优化,为政策训练提供指导轨迹样本。接着,利用所学习的策略来更新轨迹优化结果,使得引导轨迹可由当前策略重现。设计了一种学习算法,在分布式轨迹优化和策略优化之间交替进行,最终收敛到一个具有良好长期性能的解。在一个多机器人交会问题中,我们证明了该方法的有效性。在机器人仿真器上的仿真结果表明,该方法能够以较少的训练样本有效地学习分散控制策略。
Multi-robot learning has been extensively studied recently. Developing provably-correct algorithms for learning decentralized control policies remains challenging. In this letter, we propose a sample-efficient multi-robot learning method based on guided policy search to learn decentralized swarm control policies. The proposed method uses distributed trajectory optimization to provide guiding trajectory samples for policy training. In turn, the learned policy is exploited to update the trajectory optimization results so that the guiding trajectories are reproducible by the current policy. A learning algorithm is designed to alternate between distributed trajectory optimization and policy optimization, which eventually converges to a solution with good long-term performance. We demonstrate the effectiveness of our method in a multi-robot rendezvous problem. The simulation results in a robot simulator show that our method efficiently learn decentralized control policy with substantially less training samples.