Model-free Q-learning designs for linear discrete-time zero-sum games with application to H-infinity control

Model-free Q-learning designs for linear discrete-time zero-sum games with application to H-infinity control
复制标题

DOI:
10.1016/j.automatica.2006.09.019
复制
发表时间:
2007-03-01
期刊:
影响因子:
6.4
通讯作者:
Abu-Khalaf, Murad
Abu-Khalaf, Murad
中科院分区:
计算机科学2区
文献类型:
--
作者:
Al-Tamimi, Asma;Lewis, Frank L.;Abu-Khalaf, Murad

文献摘要

被引文献

相似文献

本文在不知道系统动态矩阵的情况下,在正向时间内求解与 H 无穷最优控制问题相关的离散时间线性系统二次零和博弈的最优策略。这个想法是求解零和博弈的动作相关值函数 Q(x, u, w),而不是求解满足相应博弈代数 Riccati 方程 (GARE) 的状态相关值函数 V(x)。由于状态和动作空间是连续的,因此使用两个动作网络和一个批评者网络,它们使用自适应批评者方法在前向时间上进行自适应调整。其结果是一种 Q 学习近似动态规划 (ADP) 无模型方法,可以及时解决零和博弈问题。结果表明,批评家收敛于博弈价值函数,动作网络收敛于博弈的纳什均衡。显示了算法的收敛性证明。证明该算法最终是一种求解线性二次离散时间零和博弈GARE的无模型迭代算法。通过对 F-16 飞机进行 H-infinity 控制自动驾驶仪设计,证明了该方法的有效性。 (C) 2007 Elsevier Ltd. 保留所有权利。
In this paper, the optimal strategies for discrete-time linear system quadratic zero-sum games related to the H-infinity optimal control problem are solved in forward time without knowing the system dynamical matrices. The idea is to solve for an action dependent value function Q(x, u, w) of the zero-sum game instead of solving for the state dependent value function V(x) which satisfies a corresponding game algebraic Riccati equation (GARE). Since the state and actions spaces are continuous, two action networks and one critic network are used that are adaptively tuned in forward time using adaptive critic methods. The result is a Q-learning approximate dynamic programming (ADP) model-free approach that solves the zero-sum game forward in time. It is shown that the critic converges to the game value function and the action networks converge to the Nash equilibrium of the game. Proofs of convergence of the algorithm are shown. It is proven that the algorithm ends up to be a model-free iterative algorithm to solve the GARE of the linear quadratic discrete-time zero-sum game. The effectiveness of this method is shown by performing an H-infinity control autopilot design for an F-16 aircraft. (C) 2007 Elsevier Ltd. All rights reserved.