Decentralized Adaptive Tracking Control For Large-Scale Multi-Agent Systems Under Unstructured Environment
Decentralized Adaptive Tracking Control For Large-Scale Multi-Agent Systems Under Unstructured Environment
复制标题
DOI:
10.1109/ssci51031.2022.10022141
复制
发表时间:
2022-12
期刊:
影响因子:
--
通讯作者:
Shawon Dey;Hao Xu
中科院分区:
文献类型:
--
作者:
Shawon Dey;Hao Xu
In this paper, a decentralized optimal tracking control problem has been investigated for large scale multi-agent system (LS-MAS) under unstructured environment. Due to the “Curse of Dimensionality” from a large amount of agents and constraints from the unstructured environment, conventional optimal tracking control as well as emerging mean field game and machine learning based design cannot be utilized directly. To overcome those challenges, a novel barrier function has been designed to transform the unstructured environment into a structured environment so that mean field game theory can be used to formulate decentralized optimal tracking control for LS-MAS. Then, the actor-critic-mass reinforcement learning algorithm has been developed to learn the mean field game based optimal solution under structured environment. Specifically, individual agent has three neural networks (NN), i.e., 1) mass NN that learns the behaviors of large population via estimating the solution of Fokker-Planck-Kolmogorov (FPK) equation, 2) critic NN that obtains optimal cost function by learning the solution of the Hamilton-Jacobi-Bellman (HJB) equation, 3) actor NN that solve the decentralized optimal tracking control based on the information provided by the mass and critic NN. Next, the learned decentralized optimal tracking control can be transformed from structured environment back to unstructured environment and implemented in real-time through barrier function. Overall, this developed algorithm is named MFG-based barrier-actor-critic-mass learning. The Lyapunov theorem has been used to prove the stability of the closed-loop system. Eventually, a series of numerical simulation has been conducted to demonstrate the effectiveness of the developed scheme.