Meta Proximal Policy Optimization for Cooperative Multi-Agent Continuous Control

Meta Proximal Policy Optimization for Cooperative Multi-Agent Continuous Control
复制标题

DOI:
10.1109/ijcnn55064.2022.9892004
复制
发表时间:
2022-07
期刊:
2022 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Boli Fang;Zhenghao Peng;Hao Sun;Qin Zhang
Boli Fang;Zhenghao Peng;Hao Sun;Qin Zhang
中科院分区:
其他
文献类型:
--
作者:
Boli Fang;Zhenghao Peng;Hao Sun;Qin Zhang

文献摘要

相似文献

在本文中,我们提出了多智能体代理最接近策略优化(MA 3 PO),这是一种新型的多智能体深度强化学习算法,可以解决协作连续多智能体控制的挑战。我们的方法是由大多数现有的多智能体强化学习算法主要集中在离散状态/动作空间,因此在扩展到连续状态/动作空间的环境时,计算是不可行的。为了解决计算复杂性的问题,并更好地模拟代理内部的协作,我们利用最近成功的邻近策略优化算法,有效地探索连续的动作空间,并通过元梯度方法将内在动机的概念,以刺激个体代理的行为在合作的多代理设置。为此,我们设计了代理奖励来量化个体代理级内在动机对团队级奖励的影响,并应用元梯度方法来利用这种增加,以便我们的算法可以有效地学习团队级累积奖励。在各种具有连续动作空间的多智能体强化学习基准环境上的实验表明,我们的算法不仅与现有的最先进的基准相当,而且显着降低了训练时间复杂度。
In this paper we propose Multi-Agent Proxy Proximal Policy Optimization (MA3PO), a novel multi-agent deep reinforcement learning algorithm that tackles the challenge of cooperative continuous multi-agent control. Our method is driven by the observation that most existing multi-agent reinforcement learning algorithms mainly focus on discrete state/action spaces and are thus computationally infeasible when extended to environments with continuous state/action spaces. To address the issue of computational complexity and to better model intra-agent collaboration, we make use of the recently successful Proximal Policy Optimization algorithm that effectively explores of continuous action spaces, and incorporate the notion of intrinsic motivation via meta-gradient methods so as to stimulate the behavior of individual agents in cooperative multi-agent settings. Towards these ends, we design proxy rewards to quantify the effect of individual agent-level intrinsic motivation onto the team-level reward, and apply meta-gradient methods to leverage such an addition so that our algorithm can learn the team-level cumulative reward effectively. Experiments on various multi-agent reinforcement learning benchmark environments with continuous action spaces demonstrate that our algorithm is not only comparable with the existing state-of-the-art benchmarks, but also significantly reduces training time complexity.