Online Adaptive Policy Learning Algorithm for $H_{\infty }$ State Feedback Control of Unknown Affine Nonlinear Discrete-Time Systems

Online Adaptive Policy Learning Algorithm for $H_{\infty }$ State Feedback Control of Unknown Affine Nonlinear Discrete-Time Systems
复制标题

DOI:
10.1109/tcyb.2014.2313915
复制
发表时间:
2014-07
影响因子:
11.8
通讯作者:
Huaguang Zhang;C. Qin;B. Jiang;Yanhong Luo
Huaguang Zhang;C. Qin;B. Jiang;Yanhong Luo
中科院分区:
计算机科学1区
文献类型:
--
作者:
Huaguang Zhang;C. Qin;B. Jiang;Yanhong Luo

文献摘要

被引文献

相似文献

研究了动力学未知的仿射非线性离散时间系统的H∞状态反馈控制问题。提出了一种基于自适应动态规划(ADP)的在线自适应策略学习算法(APLA),用于实时学习H∞控制问题中出现的Hamilton-Jacobi-Isaacs(HJI)方程的解。在所提出的算法中,利用三个神经网络(NN)来寻找最优值函数以及鞍点反馈控制和扰动策略的合适近似值。给出了新颖的权重更新法则,通过使用沿系统轨迹实时生成的数据来同时调整批评者、参与者和干扰神经网络。考虑到NN逼近误差,我们用Lyapunov方法对所提出的算法进行了稳定性分析。此外,通过使用神经网络识别方案,缓解了所提出算法对系统输入动态的需求。最后,仿真算例说明了所提算法的有效性。
The problem of H∞ state feedback control of affine nonlinear discrete-time systems with unknown dynamics is investigated in this paper. An online adaptive policy learning algorithm (APLA) based on adaptive dynamic programming (ADP) is proposed for learning in real-time the solution to the Hamilton-Jacobi-Isaacs (HJI) equation, which appears in the H∞ control problem. In the proposed algorithm, three neural networks (NNs) are utilized to find suitable approximations of the optimal value function and the saddle point feedback control and disturbance policies. Novel weight updating laws are given to tune the critic, actor, and disturbance NNs simultaneously by using data generated in real-time along the system trajectories. Considering NN approximation errors, we provide the stability analysis of the proposed algorithm with Lyapunov approach. Moreover, the need of the system input dynamics for the proposed algorithm is relaxed by using a NN identification scheme. Finally, simulation examples show the effectiveness of the proposed algorithm.