Decentralized Adaptive Tracking Control For Large-Scale Multi-Agent Systems Under Unstructured Environment

Decentralized Adaptive Tracking Control For Large-Scale Multi-Agent Systems Under Unstructured Environment
复制标题

DOI:
10.1109/ssci51031.2022.10022141
复制
发表时间:
2022-12
期刊:
2022 IEEE Symposium Series on Computational Intelligence (SSCI)
影响因子:
--
通讯作者:
Shawon Dey;Hao Xu
Shawon Dey;Hao Xu
中科院分区:
其他
文献类型:
--
作者:
Shawon Dey;Hao Xu

文献摘要

相似文献

研究了非结构化环境下大规模多智能体系统(LS-MAS)的分散最优跟踪控制问题。由于来自非结构化环境的大量智能体和约束的“维度诅咒”,传统的最优跟踪控制以及新兴的平均场博弈和基于机器学习的设计无法直接利用。为了克服这些挑战,设计了一种新的障碍函数,将非结构化环境转化为结构化环境,从而可以使用平均场博弈论来制定LS-MAS的分散最优跟踪控制。然后,提出了基于参与者临界质量的强化学习算法,学习结构化环境下基于平均场博弈的最优解。具体来说,个体智能体有三个神经网络,即1)质量神经网络通过估计Fokker-Planck-Kolmogorov (FPK)方程的解来学习大群体的行为,2)评论家神经网络通过学习Hamilton-Jacobi-Bellman (HJB)方程的解来获得最优成本函数,3)行动者神经网络根据质量和评论家神经网络提供的信息来求解分散最优跟踪控制。其次,将学习到的分散最优跟踪控制从结构化环境转换回非结构化环境,并通过障碍函数实时实现。总的来说,该算法被命名为基于mfg的障碍-行为体-临界-质量学习。利用李亚普诺夫定理证明了闭环系统的稳定性。最后进行了一系列数值模拟,验证了所提方案的有效性。
In this paper, a decentralized optimal tracking control problem has been investigated for large scale multi-agent system (LS-MAS) under unstructured environment. Due to the “Curse of Dimensionality” from a large amount of agents and constraints from the unstructured environment, conventional optimal tracking control as well as emerging mean field game and machine learning based design cannot be utilized directly. To overcome those challenges, a novel barrier function has been designed to transform the unstructured environment into a structured environment so that mean field game theory can be used to formulate decentralized optimal tracking control for LS-MAS. Then, the actor-critic-mass reinforcement learning algorithm has been developed to learn the mean field game based optimal solution under structured environment. Specifically, individual agent has three neural networks (NN), i.e., 1) mass NN that learns the behaviors of large population via estimating the solution of Fokker-Planck-Kolmogorov (FPK) equation, 2) critic NN that obtains optimal cost function by learning the solution of the Hamilton-Jacobi-Bellman (HJB) equation, 3) actor NN that solve the decentralized optimal tracking control based on the information provided by the mass and critic NN. Next, the learned decentralized optimal tracking control can be transformed from structured environment back to unstructured environment and implemented in real-time through barrier function. Overall, this developed algorithm is named MFG-based barrier-actor-critic-mass learning. The Lyapunov theorem has been used to prove the stability of the closed-loop system. Eventually, a series of numerical simulation has been conducted to demonstrate the effectiveness of the developed scheme.