Swarm Intelligence in Cooperative Environments: n-Step Dynamic Tree Search Algorithm Overview

Swarm Intelligence in Cooperative Environments: n-Step Dynamic Tree Search Algorithm Overview
复制标题

DOI:
10.2514/1.i011086
复制
发表时间:
2023-05
期刊:
J. Aerosp. Inf. Syst.
影响因子:
--
通讯作者:
Marc Espinós Longa;A. Tsourdos;Gokhan Inalhan
Marc Espinós Longa;A. Tsourdos;Gokhan Inalhan
中科院分区:
其他
文献类型:
--
作者:
Marc Espinós Longa;A. Tsourdos;Gokhan Inalhan

文献摘要

相似文献

强化学习基于树的规划方法在过去几年中越来越受欢迎,这是因为它们在单智能体领域取得了成功,其中有一个完美的模拟器模型:例如,围棋和国际象棋战略棋盘游戏。本文假装扩展树搜索算法的多智能体设置在一个分散的结构,处理可扩展性问题和指数增长的计算资源。[公式:动态树搜索结合了前向规划和直接时间差更新,明显优于传统的表格算法,如[公式:见正文]学习和状态-动作-奖励-状态-动作(SARSA)。未来的状态转换和奖励的预测与代理和环境之间的真实的相互作用建立和学习的模型。本文分析了随机智能逃避者的狩猎-追击合作对策的改进算法。[公式:动态树搜索旨在使单代理树搜索学习方法适应多代理边界,并且与传统的时间差技术相比,动态树搜索被证明是显著的进步。
Reinforcement learning tree-based planning methods have been gaining popularity in the last few years due to their success in single-agent domains, where a perfect simulator model is available: for example, Go and chess strategic board games. This paper pretends to extend tree search algorithms to the multiagent setting in a decentralized structure, dealing with scalability issues and exponential growth of computational resources. The [Formula: see text] dynamic tree search combines forward planning and direct temporal-difference updates, outperforming markedly conventional tabular algorithms such as [Formula: see text] learning and state-action-reward-state-action (SARSA). Future state transitions and rewards are predicted with a model built and learned from real interactions between agents and the environment. This paper analyzes the developed algorithm in the hunter–pursuit cooperative game against stochastic and intelligent evaders. The [Formula: see text] dynamic tree search aims to adapt single-agent tree search learning methods to the multiagent boundaries and is demonstrated to be a remarkable advance as compared to conventional temporal-difference techniques.