MAC-PO: Multi-Agent Experience Replay via Collective Priority Optimization

MAC-PO: Multi-Agent Experience Replay via Collective Priority Optimization
复制标题

DOI:
10.48550/arxiv.2302.10418
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Yongsheng Mei;Hanhan Zhou;Tian Lan;Guru Venkataramani;Peng Wei
Yongsheng Mei;Hanhan Zhou;Tian Lan;Guru Venkataramani;Peng Wei
中科院分区:
其他
文献类型:
--
作者:
Yongsheng Mei;Hanhan Zhou;Tian Lan;Guru Venkataramani;Peng Wei

文献摘要

相似文献

经验重放对于非策略强化学习(RL)方法至关重要。通过记忆和重用过去不同策略的经验,经验重放显著提高了RL算法的训练效率和稳定性。实践中的许多决策问题涉及多个智能体,需要集中训练分散执行范式下的多智能体强化学习。然而,现有的MARL算法往往采用标准的经验重放的过渡均匀采样,而不管他们的重要性。寻找针对MARL体验重播优化的优先采样权重还有待探索。为此,我们提出了MAC-PO,它制定了最佳的优先经验重放多智能体问题作为一个遗憾最小化的采样权重的过渡。这种优化是放松和解决使用拉格朗日乘子的方法,以获得封闭形式的最佳采样权重。通过最小化产生的政策遗憾,我们可以缩小差距之间的当前政策和名义上的最优政策,从而获得一个改进的多智能体任务的优先级计划。我们在Predator-Prey和StarCraft Multi-Agent Challenge环境中的实验结果证明了我们方法的有效性,具有更好的重播重要转换的能力,并且优于其他最先进的基线。
Experience replay is crucial for off-policy reinforcement learning (RL) methods. By remembering and reusing the experiences from past different policies, experience replay significantly improves the training efficiency and stability of RL algorithms. Many decision-making problems in practice naturally involve multiple agents and require multi-agent reinforcement learning (MARL) under centralized training decentralized execution paradigm. Nevertheless, existing MARL algorithms often adopt standard experience replay where the transitions are uniformly sampled regardless of their importance. Finding prioritized sampling weights that are optimized for MARL experience replay has yet to be explored. To this end, we propose MAC-PO, which formulates optimal prioritized experience replay for multi-agent problems as a regret minimization over the sampling weights of transitions. Such optimization is relaxed and solved using the Lagrangian multiplier approach to obtain the close-form optimal sampling weights. By minimizing the resulting policy regret, we can narrow the gap between the current policy and a nominal optimal policy, thus acquiring an improved prioritization scheme for multi-agent tasks. Our experimental results on Predator-Prey and StarCraft Multi-Agent Challenge environments demonstrate the effectiveness of our method, having a better ability to replay important transitions and outperforming other state-of-the-art baselines.