Neural network approach to continuous-time direct adaptive optimal control for partially unknown nonlinear systems

Neural network approach to continuous-time direct adaptive optimal control for partially unknown nonlinear systems
复制标题

DOI:
10.1016/j.neunet.2009.03.008
复制
发表时间:
2009-04-01
期刊:
影响因子:
7.8
通讯作者:
Lewis, Frank
Lewis, Frank
中科院分区:
计算机科学1区
文献类型:
--
作者:
Vrabie, Draguna;Lewis, Frank

文献摘要

被引文献

相似文献

在本文中,我们提出了一种在线的方法,直接自适应最优控制的无限时域成本的非线性系统在连续时间框架。该算法在线收敛到最优控制解,无需了解系统内部动态。闭环动态稳定性始终得到保证。该算法是基于油强化学习计划,即政策迭代,并利用神经网络,在一个演员/评论家的结构,参数化表示的控制策略和控制系统的性能。这两个神经网络被训练来表达最优控制器和描述无限时域控制性能的最优代价函数。算法的收敛性证明了现实的假设下,这两个神经网络不提供完美的非线性控制和成本函数的表示。其结果是一个混合控制结构,其中包括一个连续时间控制器和一个监督自适应结构,其操作的基础上从工厂和连续时间的性能动态采样的数据。这样的控制结构不同于以前在文献中看到的任何标准形式的控制器。仿真结果,考虑两个二阶非线性系统,提供。(C)2009爱思唯尔有限公司保留所有权利。
In this paper we present in a continuous-time framework an online approach to direct adaptive optimal control with infinite horizon cost for nonlinear systems. The algorithm converges online to the optimal control solution without knowledge of the internal system dynamics. Closed-loop dynamic stability is guaranteed throughout. The algorithm is based oil a reinforcement learning scheme, namely Policy iterations, and makes use of neural networks, in an Actor/Critic structure, to parametrically represent the control policy and the performance of the control system. The two neural networks are trained to express the optimal controller and optimal cost function which describes the infinite horizon control performance. Convergence of the algorithm is proven under the realistic assumption that the two neural networks do not provide perfect representations for the nonlinear control and cost functions. The result is a hybrid control structure which involves a continuous-time controller and a Supervisory adaptation structure which operates based on data sampled from the plant and from the continuous-time performance dynamics. Such control structure is unlike any standard form of controllers previously seen in the literature. Simulation results, obtained considering two second-order nonlinear systems, are provided. (C) 2009 Elsevier Ltd. All rights reserved.