Finite-Approximation-Error-Based Discrete-Time Iterative Adaptive Dynamic Programming

Finite-Approximation-Error-Based Discrete-Time Iterative Adaptive Dynamic Programming
复制标题

基于有限逼近误差的离散时间迭代自适应动态规划

DOI:
10.1109/tcyb.2014.2354377
复制
发表时间:
2014-09
影响因子:
11.8
通讯作者:
Yang, Xiong
Yang, Xiong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wei, Qinglai;Wang, Fei-Yue;Liu, Derong;Yang, Xiong

文献摘要

参考文献

被引文献

相似文献

本文提出了一种新的迭代自适应动态规划(ADP)算法,用于求解具有有限逼近误差的无限时域离散非线性系统的最优控制问题。首先,提出了一种新的ADP广义值迭代算法,使迭代性能指标函数收敛于Hamilton-Jacobi-Bellman方程的解。广义值迭代算法允许任意半正定函数初始化,克服了传统值迭代算法的缺点。在每次迭代的迭代控制律和迭代性能指标函数不能精确获得的情况下,首次建立了基于有限逼近误差的广义值迭代算法的“收敛准则设计方法”.通过自适应地设计合适的逼近误差,使迭代性能指标函数收敛到最优性能指标函数的有限邻域内。神经网络被用来实现迭代ADP算法。最后,两个仿真例子来说明所开发的方法的性能。
In this paper, a new iterative adaptive dynamic programming (ADP) algorithm is developed to solve optimal control problems for infinite horizon discrete-time nonlinear systems with finite approximation errors. First, a new generalized value iteration algorithm of ADP is developed to make the iterative performance index function converge to the solution of the Hamilton-Jacobi-Bellman equation. The generalized value iteration algorithm permits an arbitrary positive semi-definite function to initialize it, which overcomes the disadvantage of traditional value iteration algorithms. When the iterative control law and iterative performance index function in each iteration cannot accurately be obtained, for the first time a new “design method of the convergence criteria” for the finite-approximation-error-based generalized value iteration algorithm is established. A suitable approximation error can be designed adaptively to make the iterative performance index function converge to a finite neighborhood of the optimal performance index function. Neural networks are used to implement the iterative ADP algorithm. Finally, two simulation examples are given to illustrate the performance of the developed method.
使用单网络 ADP 的连续时间非线性系统非零和微分博弈的近最优控制
DOI: 10.1109/tsmcb.2012.2203336
发表时间: 2013-02
影响因子: 11.8
作者:
Huaguang Zhang;Lili Cui;Yanhong Luo
通讯作者: Yanhong Luo
DOI: 10.1109/tsmcb.2011.2148710
发表时间: 2011-10
期刊: IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics)
影响因子: --
作者:
Takeshi Mori;S. Ishii
通讯作者: Takeshi Mori;S. Ishii
利用自适应动态规划方法对未知一般非线性系统进行数据驱动的鲁棒近似最优跟踪控制
DOI: 10.1109/tnn.2011.2168538
发表时间: 2011-12
影响因子: --
作者:
Huaguang Zhang;Lili Cui;Xin Zhang;Yanhong Luo
通讯作者: Yanhong Luo
DOI: 10.13182/nse68-a17613
发表时间: 1968-03
影响因子: 1.2
作者:
Matthew M. Peet
通讯作者: Matthew M. Peet
DOI: 10.1109/tcyb.2013.2278102
发表时间: 2014-06
影响因子: 11.8
作者:
A. Jennings;R. Ordóñez
通讯作者: A. Jennings;R. Ordóñez