A deep multi-agent reinforcement learning framework for autonomous aerial navigation to grasping points on loads

A deep multi-agent reinforcement learning framework for autonomous aerial navigation to grasping points on loads
复制标题

DOI:
10.1016/j.robot.2023.104489
复制
发表时间:
2023-07
期刊:
Robotics Auton. Syst.
影响因子:
--
通讯作者:
Jingyu Chen;Ruidong Ma;J. Oyekan
Jingyu Chen;Ruidong Ma;J. Oyekan
中科院分区:
其他
文献类型:
--
作者:
Jingyu Chen;Ruidong Ma;J. Oyekan

文献摘要

相似文献

深度强化学习利用神经网络的优势,在机器人的连续控制方面取得了长足的进步。然而,在需要多个机器人相互协作完成任务的场景中,由于复杂性的增加,构建一个高效、可扩展的多智能体控制系统仍然是一个挑战。本文将每个无人机及其操纵器视为一个智能体,利用多智能体深度确定性策略梯度(madpg)的力量进行负载的协同导航和操纵。针对有针对性和灵活的场景下的导航抓点问题提出了解决方案,重点研究了如何在不依赖轨迹规划器的情况下为无人机制定无模型策略。为了克服在抓取点数量不断增加的场景中学习的挑战,我们将最优互反碰撞避免(ORCA)算法的演示纳入我们的框架中,以指导策略训练,并将两种新技术应用于MADDPG的架构中。此外,利用注意机制的课程学习,从更少的抓取点中重用知识,以方便具有更多点的负载的训练。我们的实验分别在oppeliasimsimulator上进行了3个、4个和6个抓取点负载的验证,然后使用crazyfliedadrotors将其转移到现实世界中。结果表明,在一些实际实验中,无人机从理想抓取点到最终位置的平均跟踪偏差可以小于10 cm。结果表明,与目前最先进的无模型强化学习和群优化算法相比,我们提出的方法以合理的成功率优于其他基线,特别是在抓取点较多的场景下。此外,学习到的最优策略使无人机能够到达并悬停在操作前的所有抓取点上,而不会发生碰撞。我们对定向导航和灵活导航进行了综合分析,突出了各自的优缺点。
Deep reinforcement learning, by taking advantage of neural networks, has made great strides in the continuous control of robots. However, in scenarios where multiple robots are required to collaborate with each other to accomplish a task, it is still challenging to build an efficient and scalable multi-agent control system due to increasing complexity. In this paper, we regard each unmanned aerial vehicle (UAV) with its manipulator as one agent, and leverage the power of multi-agent deep deterministic policy gradient (MADDPG) for the cooperative navigation and manipulation of a load. We propose solutions for addressing navigation to grasping point problem in targeted and flexible scenarios, and mainly focus on how to develop model-free policies for the UAVs without relying on a trajectory planner. To overcome the challenges of learning in scenarios with an increasing number of grasping points, we incorporate the demonstrations from an Optimal Reciprocal Collision Avoidance (ORCA) algorithm into our framework to guide the policy training and adapt two novel techniques into the architecture of MADDPG. Furthermore, curriculum learning with the attention mechanism is utilized by reusing knowledge from fewer grasping points to facilitate the training of a load with more points. Our experiments were validated by a load with three, four and six grasping points respectively inCoppeliasimsimulator and then transferred into the real world withCrazyfliequadrotors. Our results show that the average tracking deviations from the desirable grasping point to the final position of the UAV can be less than 10 cm in some real-world experiments. Compared with state-of-the-art model-free reinforcement learning and swarm optimization algorithms, results show that our proposed methods outperform other baselines with a reasonable success rate especially in the scenarios with more grasping points. Furthermore, the learned optimal policies enable UAVs to reach and hover over all the grasping points before manipulation without any collision. We conducted a comprehensive analysis of both targeted and flexible navigation, highlighting their respective advantages and disadvantages.