Local Advantage Actor-Critic for Robust Multi-Agent Deep Reinforcement Learning

Local Advantage Actor-Critic for Robust Multi-Agent Deep Reinforcement Learning
复制标题

DOI:
10.1109/mrs50823.2021.9620607
复制
发表时间:
2021-10
期刊:
2021 International Symposium on Multi-Robot and Multi-Agent Systems (MRS)
影响因子:
--
通讯作者:
Yuchen Xiao;Xueguang Lyu;Chris Amato
Yuchen Xiao;Xueguang Lyu;Chris Amato
中科院分区:
其他
文献类型:
--
作者:
Yuchen Xiao;Xueguang Lyu;Chris Amato

文献摘要

相似文献

策略梯度方法在多代理强化学习中已经变得流行,但是由于环境随机性和探索代理的存在,它们遭受高方差(即,非平稳性),这可能由于信用分配的困难而恶化。因此,需要一种不仅能够有效地解决上述两个问题,而且足够鲁棒以解决各种任务的方法。为此,我们提出了一种新的多智能体策略梯度方法,称为鲁棒局部优势(ROLA)Actor-Critic。ROLA允许每个代理学习一个单独的行动价值函数作为本地的评论家,以及改善环境的非平稳性通过一种新的集中式训练方法的基础上集中的评论家。通过使用这个本地评论家,每个代理计算一个基线,以减少其政策梯度估计的方差,这导致了一个预期的优势行动价值超过其他代理的选择,隐含地改善信用分配。我们评估ROLA在不同的基准,并显示其鲁棒性和有效性超过了一些国家的最先进的多智能体的政策梯度算法。
Policy gradient methods have become popular in multi-agent reinforcement learning, but they suffer from high variance due to the presence of environmental stochasticity and exploring agents (i.e., non-stationarity), which is potentially worsened by the difficulty in credit assignment. As a result, there is a need for a method that is not only capable of efficiently solving the above two problems but also robust enough to solve a variety of tasks. To this end, we propose a new multi-agent policy gradient method, called Robust Local Advantage (ROLA) Actor-Critic. ROLA allows each agent to learn an individual action-value function as a local critic as well as ameliorating environment non-stationarity via a novel centralized training approach based on a centralized critic. By using this local critic, each agent calculates a baseline to reduce variance on its policy gradient estimation, which results in an expected advantage action-value over other agents' choices that implicitly improves credit assignment. We evaluate ROLA across diverse benchmarks and show its robustness and effectiveness over a number of state-of-the-art multi-agent policy gradient algorithms.