Hamiltonian-Driven Adaptive Dynamic Programming for Continuous Nonlinear Dynamical Systems

Hamiltonian-Driven Adaptive Dynamic Programming for Continuous Nonlinear Dynamical Systems
复制标题

DOI:
10.1109/tnnls.2017.2654324
复制
发表时间:
2017-02
影响因子:
10.4
通讯作者:
Yongliang Yang;D. Wunsch;Yixin Yin
Yongliang Yang;D. Wunsch;Yixin Yin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yongliang Yang;D. Wunsch;Yixin Yin

文献摘要

被引文献

相似文献

针对连续时间非线性系统,提出了一种Hamilton驱动的自适应动态规划(ADP)框架,包括对容许控制的评价、两种不同容许策略相对于相应性能函数的比较以及容许控制的性能改进.结果表明,哈密顿量可以作为连续时间系统的时间差分。在Hamilton驱动的ADP中,训练评价器网络以输出值梯度。然后,评论家和系统动力学之间的内积产生价值导数。在一定条件下,哈密尔顿泛函的极小化等价于值函数逼近。从任意容许控制出发,给出了最优控制逼近的迭代算法,并证明了其收敛性。实现是通过神经网络近似。两个仿真研究证明了哈密顿驱动的ADP的有效性。
This paper presents a Hamiltonian-driven framework of adaptive dynamic programming (ADP) for continuous time nonlinear systems, which consists of evaluation of an admissible control, comparison between two different admissible policies with respect to the corresponding the performance function, and the performance improvement of an admissible control. It is showed that the Hamiltonian can serve as the temporal difference for continuous-time systems. In the Hamiltonian-driven ADP, the critic network is trained to output the value gradient. Then, the inner product between the critic and the system dynamics produces the value derivative. Under some conditions, the minimization of the Hamiltonian functional is equivalent to the value function approximation. An iterative algorithm starting from an arbitrary admissible control is presented for the optimal control approximation with its convergence proof. The implementation is accomplished by a neural network approximation. Two simulation studies demonstrate the effectiveness of Hamiltonian-driven ADP.