Policy Iteration Adaptive Dynamic Programming Algorithm for Discrete-Time Nonlinear Systems

Policy Iteration Adaptive Dynamic Programming Algorithm for Discrete-Time Nonlinear Systems
复制标题

离散时间非线性系统的策略迭代自适应动态规划算法

DOI:
10.1109/tnnls.2013.2281663
复制
发表时间:
2014-03-01
影响因子:
10.4
通讯作者:
Wei, Qinglai
Wei, Qinglai
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liu, Derong;Wei, Qinglai

文献摘要

被引文献

相似文献

提出了一种新的离散时间策略迭代自适应动态规划(ADP)方法,用于求解非线性系统的无限时域最优控制问题。其思想是使用迭代ADP技术获得迭代控制律,优化迭代性能指标函数。本文的主要贡献是首次分析了离散时间非线性系统策略迭代方法的收敛性和稳定性。证明了迭代性能指标函数非增收敛于Hamilton-Jacobi-Bellman方程的最优解。证明了任意迭代控制律都能镇定非线性系统。神经网络分别用于逼近性能指标函数和计算最优控制律,以便于迭代ADP算法的实现,其中的权矩阵的收敛性进行了分析。最后,数值结果和分析,以说明所开发的方法的性能。
This paper is concerned with a new discrete-time policy iteration adaptive dynamic programming (ADP) method for solving the infinite horizon optimal control problem of nonlinear systems. The idea is to use an iterative ADP technique to obtain the iterative control law, which optimizes the iterative performance index function. The main contribution of this paper is to analyze the convergence and stability properties of policy iteration method for discrete-time nonlinear systems for the first time. It shows that the iterative performance index function is nonincreasingly convergent to the optimal solution of the Hamilton-Jacobi-Bellman equation. It is also proven that any of the iterative control laws can stabilize the nonlinear systems. Neural networks are used to approximate the performance index function and compute the optimal control law, respectively, for facilitating the implementation of the iterative ADP algorithm, where the convergence of the weight matrices is analyzed. Finally, the numerical results and analysis are presented to illustrate the performance of the developed method.