A Deeper Understanding of State-Based Critics in Multi-Agent Reinforcement Learning

A Deeper Understanding of State-Based Critics in Multi-Agent Reinforcement Learning
复制标题

DOI:
10.1609/aaai.v36i9.21171
复制
发表时间:
2022-01
期刊:
--
影响因子:
--
通讯作者:
Xueguang Lyu;Andrea Baisero;Yuchen Xiao;Chris Amato
Xueguang Lyu;Andrea Baisero;Yuchen Xiao;Chris Amato
中科院分区:
其他
文献类型:
--
作者:
Xueguang Lyu;Andrea Baisero;Yuchen Xiao;Chris Amato

文献摘要

相似文献

分散执行的集中式训练,即以集中式离线方式进行训练,已成为多智能体强化学习中流行的解决方案范例。许多这样的方法采用基于状态的批评者的行动者-批评者的形式,因为集中式训练允许访问真实的系统状态,这在训练期间可能是有用的,尽管在执行时不可用。基于国家的批评已成为一种常见的经验选择,尽管这种选择的理论依据或分析有限。在本文中,我们表明,基于状态的批评可以引入偏见的政策梯度估计,潜在地破坏了算法的渐近保证。我们还表明,即使基于状态的批评不引入任何偏见,他们仍然可以导致一个更大的梯度方差,与共同的直觉。最后,我们展示了理论在实践中的影响,通过比较不同形式的集中批评在广泛的共同基准,并详细说明各种环境属性是如何与不同类型的批评的有效性。
Centralized Training for Decentralized Execution, where training is done in a centralized offline fashion, has become a popular solution paradigm in Multi-Agent Reinforcement Learning. Many such methods take the form of actor-critic with state-based critics, since centralized training allows access to the true system state, which can be useful during training despite not being available at execution time. State-based critics have become a common empirical choice, albeit one which has had limited theoretical justification or analysis. In this paper, we show that state-based critics can introduce bias in the policy gradient estimates, potentially undermining the asymptotic guarantees of the algorithm. We also show that, even if the state-based critics do not introduce any bias, they can still result in a larger gradient variance, contrary to the common intuition. Finally, we show the effects of the theories in practice by comparing different forms of centralized critics on a wide range of common benchmarks, and detail how various environmental properties are related to the effectiveness of different types of critics.