Contrasting Centralized and Decentralized Critics in Multi-Agent Reinforcement Learning

Contrasting Centralized and Decentralized Critics in Multi-Agent Reinforcement Learning
复制标题

DOI:
10.5555/3463952.3464053
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Xueguang Lyu;Yuchen Xiao;Brett Daley;Chris Amato
Xueguang Lyu;Yuchen Xiao;Brett Daley;Chris Amato
中科院分区:
其他
文献类型:
--
作者:
Xueguang Lyu;Yuchen Xiao;Brett Daley;Chris Amato

文献摘要

被引文献

相似文献

分散式执行的集中式训练(Centralized Training for Decentralized Execution),即使用集中式信息离线训练智能体,但在线上以分散的方式执行,在多智能体强化学习社区中得到了普及。特别是,集中的评论家和分散的演员的演员-评论家方法是这种想法的一个常见实例。然而,在这种情况下使用集中式批评的含义并没有得到充分的讨论和理解,即使它是许多算法的标准选择。因此,我们正式分析了集中式和分散式的批评方法,从而更深入地了解了批评选择的含义。因为我们的理论做出了不切实际的假设,我们还在广泛的环境中对集中式和分散式批评方法进行了实证比较,以验证我们的理论并提供实用的建议。我们表明,目前文献中存在关于集中式批评的误解,并表明集中式批评设计并不是严格有益的,而是集中式和分散式批评都有不同的优点和缺点,算法设计者应该考虑到这些优点和缺点。
Centralized Training for Decentralized Execution, where agents are trained offline using centralized information but execute in a decentralized manner online, has gained popularity in the multi-agent reinforcement learning community. In particular, actor-critic methods with a centralized critic and decentralized actors are a common instance of this idea. However, the implications of using a centralized critic in this context are not fully discussed and understood even though it is the standard choice of many algorithms. We therefore formally analyze centralized and decentralized critic approaches, providing a deeper understanding of the implications of critic choice. Because our theory makes unrealistic assumptions, we also empirically compare the centralized and decentralized critic methods over a wide set of environments to validate our theories and to provide practical advice. We show that there exist misconceptions regarding centralized critics in the current literature and show that the centralized critic design is not strictly beneficial, but rather both centralized and decentralized critics have different pros and cons that should be taken into account by algorithm designers.