MAVEN: Multi-Agent Variational Exploration

MAVEN: Multi-Agent Variational Exploration
复制标题

DOI:
--
复制
发表时间:
2019-10
影响因子:
7.8
通讯作者:
Anuj Mahajan;Tabish Rashid;Mikayel Samvelyan;Shimon Whiteson
Anuj Mahajan;Tabish Rashid;Mikayel Samvelyan;Shimon Whiteson
中科院分区:
数学1区
文献类型:
--
作者:
Anuj Mahajan;Tabish Rashid;Mikayel Samvelyan;Shimon Whiteson

文献摘要

被引文献

相似文献

由于执行期间的通信限制和训练中的计算易处理性,具有分散执行的集中训练是协作深度多智能体强化学习的重要设置。在本文中,我们分析了已知在复杂环境中具有卓越性能的基于价值的方法[43]。我们特别关注 QMIX [40],它是该领域当前最先进的技术。我们表明,QMIX 和类似方法引入的联合动作值的表征约束导致探索效果不佳和次优。此外,我们提出了一种名为 MAVEN 的新颖方法,通过引入分层控制的潜在空间来混合价值和基于策略的方法。基于价值的代理将其行为限制在由分层策略控制的共享潜在变量上。这使得 MAVEN 能够实现承诺的、暂时扩展的探索,这是解决复杂的多智能体任务的关键。我们的实验结果表明,MAVEN 在具有挑战性的 SMAC 域上实现了显着的性能改进 [43]。
Centralised training with decentralised execution is an important setting for cooperative deep multi-agent reinforcement learning due to communication constraints during execution and computational tractability in training. In this paper, we analyse value-based methods that are known to have superior performance in complex environments [43]. We specifically focus on QMIX [40], the current state-of-the-art in this domain. We show that the representational constraints on the joint action-values introduced by QMIX and similar methods lead to provably poor exploration and suboptimality. Furthermore, we propose a novel approach called MAVEN that hybridises value and policy-based methods by introducing a latent space for hierarchical control. The value-based agents condition their behaviour on the shared latent variable controlled by a hierarchical policy. This allows MAVEN to achieve committed, temporally extended exploration, which is key to solving complex multi-agent tasks. Our experimental results show that MAVEN achieves significant performance improvements on the challenging SMAC domain [43].