Neural-network-based online optimal control for uncertain non-linear continuous-time systems with control constraints

Neural-network-based online optimal control for uncertain non-linear continuous-time systems with control constraints
复制标题

DOI:
10.1049/iet-cta.2013.0472
复制
发表时间:
2013-11
影响因子:
2.6
通讯作者:
Xiong Yang;Derong Liu;Yuzhu Huang
Xiong Yang;Derong Liu;Yuzhu Huang
中科院分区:
计算机科学4区
文献类型:
--
作者:
Xiong Yang;Derong Liu;Yuzhu Huang

文献摘要

被引文献

相似文献

本研究针对不确定非线性连续时间系统的无限时域最优控制问题,提出一种在线自适应最优控制方法,其控制策略具有饱和约束。提出了一种新的识别-批评结构,使用两个神经网络(NN)来近似Hamilton-Jacobi-Bellman方程:一个识别NN用于估计不确定系统动态,一个批评NN用于获得最优控制,而不是典型的行动-批评双网络用于强化学习。基于所开发的架构,标识符NN和评论家NN同时调谐。同时,与政策迭代中必不可少的初始稳定控制不同,对初始控制没有特殊要求。此外,通过使用李亚普诺夫直接方法,在保持闭环系统稳定的情况下,保证了辨识器神经网络和批评器神经网络的权值一致最终有界。最后,通过一个算例验证了该方法的有效性.
In this study, an online adaptive optimal control scheme is developed for solving the infinite-horizon optimal control problem of uncertain non-linear continuous-time systems with the control policy having saturation constraints. A novel identifier-critic architecture is presented to approximate the Hamilton-Jacobi-Bellman equation using two neural networks (NNs): an identifier NN is used to estimate the uncertain system dynamics and a critic NN is utilised to derive the optimal control instead of typical action-critic dual networks employed in reinforcement learning. Based on the developed architecture, the identifier NN and the critic NN are tuned simultaneously. Meanwhile, unlike initial stabilising control indispensable in policy iteration, there is no special requirement imposed on the initial control. Moreover, by using Lyapunov's direct method, the weights of the identifier NN and the critic NN are guaranteed to be uniformly ultimately bounded, while keeping the closed-loop system stable. Finally, an example is provided to demonstrate the effectiveness of the present approach.