DARL1N: Distributed multi-Agent Reinforcement Learning with One-hop Neighbors

DARL1N: Distributed multi-Agent Reinforcement Learning with One-hop Neighbors
复制标题

DOI:
10.1109/iros47612.2022.9981441
复制
发表时间:
2022-02
期刊:
2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Baoqian Wang;Junfei Xie;Nikolay A. Atanasov
Baoqian Wang;Junfei Xie;Nikolay A. Atanasov
中科院分区:
其他
文献类型:
--
作者:
Baoqian Wang;Junfei Xie;Nikolay A. Atanasov

文献摘要

被引文献

相似文献

随着智能体数量的增加,多智能体强化学习(MARL)方法在策略和值函数表示中面临着维度的诅咒。分布式或并行训练技术的发展也受到智能体动态之间全局耦合的阻碍,这需要同时进行状态转换。介绍了带有一跳邻居的分布式多智能体强化学习(DARLIN)。DARLIN是一种脱出策略的actor-critic MARL方法,它打破了维度的诅咒,通过将智能体交互限制在单跳邻域来实现分布式训练。每个代理在一个单跳邻域上优化其值和策略功能,降低了表示复杂性,同时通过使用不同数量和状态的邻域进行训练来保持表达性。这种结构实现了DARLIN的关键贡献:一个分布式训练过程,其中每个计算节点只模拟一小部分代理的状态转换,大大加快了大规模MARL策略的训练。与最先进的MARL方法的比较表明,随着智能体数量的增加,DARLIN在不牺牲策略质量的情况下显著减少了训练时间。
Multi-agent reinforcement learning (MARL) meth-ods face a curse of dimensionality in the policy and value function representations as the number of agents increases. The development of distributed or parallel training techniques is also hindered by the global coupling among the agent dynamics, requiring simultaneous state transitions. This paper introduces Distributed multi-Agent Reinforcement Learning with One-hop Neighbors (DARLIN). DARLIN is an off-policy actor-critic MARL method that breaks the curse of dimensionality and achieves distributed training by restricting the agent interactions to one-hop neighborhoods. Each agent optimizes its value and policy functions over a one-hop neighborhood, reducing the representation complexity, yet maintaining expressiveness by training with varying numbers and states of neighbors. This structure enables the key contribution of DARLIN: a distributed training procedure in which each compute node simulates the state transitions of only a small subset of the agents, greatly accelerating the training of large-scale MARL policies. Comparisons with state-of-the-art MARL methods show that DARLIN significantly reduces training time without sacrificing policy quality as the number of agents increases.