Leader-Based Optimal Coordination Control for the Consensus Problem of Multiagent Differential Games via Fuzzy Adaptive Dynamic Programming

Leader-Based Optimal Coordination Control for the Consensus Problem of Multiagent Differential Games via Fuzzy Adaptive Dynamic Programming
复制标题

基于领导者的模糊自适应动态规划多智能体微分博弈共识问题最优协调控制

DOI:
10.1109/tfuzz.2014.2310238
复制
发表时间:
2015-02-01
影响因子:
11.9
通讯作者:
Luo, Yanhong
Luo, Yanhong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhang, Huaguang;Zhang, Jilie;Luo, Yanhong

文献摘要

被引文献

相似文献

在本文中,提出了一种新的在线方案,以通过模糊自适应动态编程为多种差异游戏的共识问题设计最佳的协调控制,该节目将游戏理论汇集在一起​​,广义模糊双曲线模型(GFHM)和自适应动态编程。一般而言,多种差异游戏的最佳协调控制是耦合的汉密尔顿 - 雅各比(HJ)方程的解决方案。在这里,基于策略迭代算法,GFHM首次使用GFHM来近似耦合HJ方程的解决方案(值函数)。也就是说,对于每个代理,GFHM用于捕获本地共识误差和本地值函数之间的映射。由于我们的方案使用每个代理的单网络体系结构(这消除了与双NETWORK架构相比的动作网络模型),因此它是多元素系统的更合理的体系结构。此外,利用近似解决方案获得最佳配位控制。最后,我们为我们的方案提供了稳定性分析,并证明权重估计误差和局部共识误差最终是界限的。此外,控制节点轨迹被证明是统一的合作轨迹。
In this paper, a new online scheme is presented to design the optimal coordination control for the consensus problem of multiagent differential games by fuzzy adaptive dynamic programming, which brings together game theory, generalized fuzzy hyperbolic model (GFHM), and adaptive dynamic programming. In general, the optimal coordination control for multiagent differential games is the solution of the coupled Hamilton-Jacobi (HJ) equations. Here, for the first time, GFHMs are used to approximate the solutions (value functions) of the coupled HJ equations, based on policy iteration algorithm. Namely, for each agent, GFHM is used to capture the mapping between the local consensus error and local value function. Since our scheme uses the single-network architecture for each agent (which eliminates the action network model compared with dual-network architecture), it is a more reasonable architecture for multiagent systems. Furthermore, the approximation solution is utilized to obtain the optimal coordination control. Finally, we give the stability analysis for our scheme, and prove the weight estimation error and the local consensus error are uniformly ultimately bounded. Further, the control node trajectory is proven to be cooperative uniformly ultimately bounded.