Swarm Intelligence in Cooperative Environments: N-Step Dynamic Tree Search Algorithm Extended Analysis

Swarm Intelligence in Cooperative Environments: N-Step Dynamic Tree Search Algorithm Extended Analysis
复制标题

DOI:
10.23919/acc53348.2022.9867171
复制
发表时间:
2022-06
期刊:
2022 American Control Conference (ACC)
影响因子:
--
通讯作者:
Marc Espinós Longa;A. Tsourdos;G. Inalhan
Marc Espinós Longa;A. Tsourdos;G. Inalhan
中科院分区:
其他
文献类型:
--
作者:
Marc Espinós Longa;A. Tsourdos;G. Inalhan

文献摘要

相似文献

在过去的几年里,基于强化学习树的规划方法由于在单智能体领域的成功而越来越受欢迎,在单智能体领域,可以获得完美的模拟器模型,例如围棋和国际象棋战略棋类游戏。本文试图将树搜索算法扩展到分布式结构中的多智能体环境,以处理可伸缩性问题和计算资源的指数增长。N步动态树搜索结合了前向规划和直接时间差分更新,显著优于Q-学习和SARSA等最先进的算法。通过建立一个模型,并从代理与环境之间的真实交互中学习,预测未来的状态转换和回报。本文在前人工作的基础上,分析了智能躲避猎人-追捕合作博弈中的改进算法。N步动态树搜索旨在将最成功的单智能体学习方法应用于多智能体边界,与传统的时差技术相比是一个显着的进步。
Reinforcement learning tree-based planning methods have been gaining popularity in the last few years due to their success in single-agent domains, where a perfect simulator model is available, e.g., Go and chess strategic board games. This paper pretends to extend tree search algorithms to the multi-agent setting in a decentralized structure, dealing with scalability issues and exponential growth of computational resources. The N-Step Dynamic Tree Search combines forward planning and direct temporal-difference updates, outperforming markedly state-of-the-art algorithms such as Q-Learning and SARSA. Future state transitions and rewards are predicted with a model built and learned from real interactions between agents and the environment. As an extension of previous work, this paper analyses the developed algorithm in the Hunter-Pursuit cooperative game against intelligent evaders. The N-Step Dynamic Tree Search aims to adapt the most successful single-agent learning methods to the multi-agent boundaries and demonstrates to be a remarkable advance compared to conventional temporal-difference techniques.