Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics

Reducing Overestimation Bias in Multi-Agent Domains Using Double Centralized Critics
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
J. Ackermann;Volker Gabler;Takayuki Osa;Masashi Sugiyama
J. Ackermann;Volker Gabler;Takayuki Osa;Masashi Sugiyama
中科院分区:
其他
文献类型:
--
作者:
J. Ackermann;Volker Gabler;Takayuki Osa;Masashi Sugiyama

文献摘要

被引文献

相似文献

许多真实的任务需要多个智能体协同工作。近年来提出了多智能体强化学习(RL)方法来解决这些任务,但目前的方法往往无法有效地学习策略。因此,我们调查存在一个共同的弱点,在单代理RL,即价值函数高估偏见,在多代理设置。根据我们的研究结果,我们提出了一种方法,通过使用双中心的批评,减少这种偏见。我们评估它的六个混合合作竞争的任务,显示出显着的优势,比目前的方法。最后,我们研究了多智能体方法在高维机器人任务中的应用,并表明我们的方法可以用来学习分散的策略。
Many real world tasks require multiple agents to work together. Multi-agent reinforcement learning (RL) methods have been proposed in recent years to solve these tasks, but current methods often fail to efficiently learn policies. We thus investigate the presence of a common weakness in single-agent RL, namely value function overestimation bias, in the multi-agent setting. Based on our findings, we propose an approach that reduces this bias by using double centralized critics. We evaluate it on six mixed cooperative-competitive tasks, showing a significant advantage over current methods. Finally, we investigate the application of multi-agent methods to high-dimensional robotic tasks and show that our approach can be used to learn decentralized policies in this domain.