Asynchronous Actor-Critic for Multi-Agent Reinforcement Learning

Asynchronous Actor-Critic for Multi-Agent Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2209.10113
复制
发表时间:
2022-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuchen Xiao;Weihao Tan;Chris Amato
Yuchen Xiao;Weihao Tan;Chris Amato
中科院分区:
其他
文献类型:
--
作者:
Yuchen Xiao;Weihao Tan;Chris Amato

文献摘要

被引文献

相似文献

在现实环境中,跨多个代理同步决策是有问题的,因为它需要代理等待其他代理终止并可靠地进行终止通信。理想情况下,代理应该异步学习和执行。这样的异步方法还允许时间上延长的动作,其可以基于所执行的情况和动作而花费不同的时间量。不幸的是,目前的策略梯度方法不适用于异步设置,因为它们假设代理同步的原因在每个时间步的动作选择。为了允许异步学习和决策,我们制定了一套异步多代理actor-critic方法,允许代理在三个标准的训练范式中直接优化异步策略:分散式学习,集中式学习和分散式执行的集中式训练。在各种现实领域的实证结果(在模拟和硬件)证明了我们的方法在大型多智能体问题的优越性,并验证了我们的算法学习高质量和异步解决方案的有效性。
Synchronizing decisions across multiple agents in realistic settings is problematic since it requires agents to wait for other agents to terminate and communicate about termination reliably. Ideally, agents should learn and execute asynchronously instead. Such asynchronous methods also allow temporally extended actions that can take different amounts of time based on the situation and action executed. Unfortunately, current policy gradient methods are not applicable in asynchronous settings, as they assume that agents synchronously reason about action selection at every time step. To allow asynchronous learning and decision-making, we formulate a set of asynchronous multi-agent actor-critic methods that allow agents to directly optimize asynchronous policies in three standard training paradigms: decentralized learning, centralized learning, and centralized training for decentralized execution. Empirical results (in simulation and hardware) in a variety of realistic domains demonstrate the superiority of our approaches in large multi-agent problems and validate the effectiveness of our algorithms for learning high-quality and asynchronous solutions.