CM3: Cooperative Multi-goal Multi-stage Multi-agent Reinforcement Learning

CM3: Cooperative Multi-goal Multi-stage Multi-agent Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2018
期刊:
--
影响因子:
--
通讯作者:
Jiachen Yang;A. Nakhaei;David Isele;Hongyuan Zha;K. Fujimura
Jiachen Yang;A. Nakhaei;David Isele;Hongyuan Zha;K. Fujimura
中科院分区:
其他
文献类型:
--
作者:
Jiachen Yang;A. Nakhaei;David Isele;Hongyuan Zha;K. Fujimura

文献摘要

相似文献

我们提出了CM3,一种新的深度强化学习方法,用于合作多智能体问题,其中智能体必须协调以共同成功地实现不同的个体目标。我们将多智能体学习重构为两阶段的课程,包括学习完成单个任务的单智能体阶段,以及学习在其他智能体存在下进行合作的多智能体阶段。这两个阶段是通过神经网络策略和价值函数的模块化增强连接起来的。我们通过制定政策梯度的地方和全球观点,并通过双重批评(由分散的价值函数和集中的行动价值函数组成)学习,进一步调整行为-批评框架以适应本课程。我们在一个新的具有稀疏奖励的高维多智能体环境中评估了CM3:在模拟城市移动(SUMO)交通模拟器中,多辆自动驾驶汽车之间协商变道。详细的消融实验表明,CM3中的每个组件都有积极的贡献,并且总体合成比现有的协作多智能体方法更快地收敛到更高性能的策略。
We propose CM3, a new deep reinforcement learning method for cooperative multi-agent problems where agents must coordinate for joint success in achieving different individual goals. We restructure multi-agent learning into a two-stage curriculum, consisting of a single-agent stage for learning to accomplish individual tasks, followed by a multi-agent stage for learning to cooperate in the presence of other agents. These two stages are bridged by modular augmentation of neural network policy and value functions. We further adapt the actor-critic framework to this curriculum by formulating local and global views of the policy gradient and learning via a double critic, consisting of a decentralized value function and a centralized action-value function. We evaluated CM3 on a new high-dimensional multi-agent environment with sparse rewards: negotiating lane changes among multiple autonomous vehicles in the Simulation of Urban Mobility (SUMO) traffic simulator. Detailed ablation experiments show the positive contribution of each component in CM3, and the overall synthesis converges significantly faster to higher performance policies than existing cooperative multi-agent methods.