QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning

QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2018-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Tabish Rashid;Mikayel Samvelyan;C. S. D. Witt;Gregory Farquhar;Jakob N. Foerster;Shimon Whiteson
Tabish Rashid;Mikayel Samvelyan;C. S. D. Witt;Gregory Farquhar;Jakob N. Foerster;Shimon Whiteson
中科院分区:
其他
文献类型:
--
作者:
Tabish Rashid;Mikayel Samvelyan;C. S. D. Witt;Gregory Farquhar;Jakob N. Foerster;Shimon Whiteson

文献摘要

被引文献

相似文献

在许多现实世界中,一组代理必须协调他们的行为,同时以分散的方式行事。与此同时,通常可以在模拟或实验室环境中以集中方式训练代理,其中可以获得全局状态信息并解除通信约束。学习以额外状态信息为条件的联合动作值是利用集中式学习的一种有吸引力的方法,但提取分散式策略的最佳策略尚不清楚。我们的解决方案是QMIX,这是一种新颖的基于价值的方法,可以以集中的端到端方式训练分散的策略。QMIX采用一个网络,该网络将联合动作值估计为每个代理值的复杂非线性组合,仅以局部观察为条件。我们在结构上强制联合行动值在每个代理值中是单调的,这允许在非策略学习中最大化联合行动值,并保证集中式和分散式策略之间的一致性。我们评估QMIX在一组具有挑战性的星际争霸II微观管理任务,并表明QMIX显着优于现有的基于值的多智能体强化学习方法。
In many real-world settings, a team of agents must coordinate their behaviour while acting in a decentralised way. At the same time, it is often possible to train the agents in a centralised fashion in a simulated or laboratory setting, where global state information is available and communication constraints are lifted. Learning joint action-values conditioned on extra state information is an attractive way to exploit centralised learning, but the best strategy for then extracting decentralised policies is unclear. Our solution is QMIX, a novel value-based method that can train decentralised policies in a centralised end-to-end fashion. QMIX employs a network that estimates joint action-values as a complex non-linear combination of per-agent values that condition only on local observations. We structurally enforce that the joint-action value is monotonic in the per-agent values, which allows tractable maximisation of the joint action-value in off-policy learning, and guarantees consistency between the centralised and decentralised policies. We evaluate QMIX on a challenging set of StarCraft II micromanagement tasks, and show that QMIX significantly outperforms existing value-based multi-agent reinforcement learning methods.