Swarm Intelligence in Cooperative Environments: Introducing the N-Step Dynamic Tree Search Algorithm

Swarm Intelligence in Cooperative Environments: Introducing the N-Step Dynamic Tree Search Algorithm
复制标题

DOI:
10.2514/6.2022-1839
复制
发表时间:
2022-01
期刊:
AIAA SCITECH 2022 Forum
影响因子:
--
通讯作者:
Marc Espinós Longa;G. Inalhan;A. Tsourdos
Marc Espinós Longa;G. Inalhan;A. Tsourdos
中科院分区:
其他
文献类型:
--
作者:
Marc Espinós Longa;G. Inalhan;A. Tsourdos

文献摘要

相似文献

有关环境动态的不确定性和部分或未知信息导致基于奖励的方法在单智能体和多智能体学习问题中发挥关键作用。基于树的规划方法(例如蒙特卡罗树搜索算法)在可以使用完美模拟器模型的单代理领域(例如围棋和国际象棋战略棋盘游戏)取得了惊人的成功。本文提出了一种基于树的分散规划方案,该方案将前向规划与应用于多智能体设置的直接强化学习时差更新相结合。前瞻性规划需要一个从经验中学习并通过函数近似表示的发动机模型。评估和验证在 Hunter-Prey Pursuit 合作环境中进行,并将性能与最先进的 RL 技术进行比较。 N 步动态树搜索(NSDTS)假装将最成功的单智能体学习方法适应去中心化系统结构中的多智能体边界,解决中心化系统所遭受的可扩展性问题和计算资源的指数增长。与传统的 Q-Learning 时差方法相比,NSDTS 被证明是一个显着的进步
Uncertainty and partial or unknown information about environment dynamics have led reward-based methods to play a key role in the Single-Agent and Multi-Agent Learning problem. Tree-based planning approaches such as Monte Carlo Tree Search algorithm have been a striking success in single-agent domains where a perfect simulator model is available, e.g., Go and chess strategic board games. This paper presents a decentralized tree-based planning scheme, that combines forward planning with direct reinforcement learning temporal-difference updates applied to the multi-agent setting. Forward planning requires an engine model which is learned from experience and represented via function approximation. Evaluation and validation are carried out in the Hunter-Prey Pursuit cooperative environment and performance is compared with state-of-the-art RL techniques. N-Step Dynamic Tree Search (NSDTS) pretends to adapt the most successful single-agent learning methods to the multi-agent boundaries in a decentralized system structure, dealing with scalability issues and exponential growth of computational resources suffered by centralized systems. NSDTS demonstrates to be a remarkable advance compared to the conventional Q-Learning temporal-difference method