An identifier-actor-optimizer policy learning architecture for optimal control of continuous-time nonlinear systems

An identifier-actor-optimizer policy learning architecture for optimal control of continuous-time nonlinear systems
复制标题

用于连续时间非线性系统最优控制的标识符-参与者-优化器策略学习架构

DOI:
10.1007/s11433-019-1481-2
复制
发表时间:
2020
期刊:
Science China-Physics,Mechanics & Astronomy
影响因子:
--
通讯作者:
Li Junfeng
Li Junfeng
中科院分区:
其他
文献类型:
--
作者:
Cheng Lin;Wang Zhenbo;Jiang Fanghua;Li Junfeng

文献摘要

被引文献

相似文献

针对连续时间非线性系统的实时最优控制问题,提出了一种基于辨识器-执行器-优化器(IAO)策略学习结构的智能求解方法。在这种基于IAO的策略学习方法中,使用深度神经网络(DNN)开发了一个动态标识符来近似系统动态的未知部分。然后,提出了一种基于间接方法的优化器,以产生高质量的最优行动的系统控制考虑的约束条件和性能指标。此外,开发了一个基于DNN的Actor来近似获得的最佳动作,并将良好的初始猜测返回给优化器。通过这种方式,传统的最优控制方法和国家的最先进的DNN技术相结合,在基于IAO的最优策略学习方法。相对于Actor-Critic结构的强化学习算法存在的奖励设计困难和计算效率低的问题,基于IAO的最优策略学习算法在求解复杂连续时间最优控制问题(OCPs)时,具有自定义参数少、学习速度快、收敛性好等优点.三个空间飞行控制任务的仿真结果证实了这种基于IAO的策略学习策略的有效性,并说明了开发的基于DNN的连续时间OCP的最优控制方法的性能。
An intelligent solution method is proposed to achieve real-time optimal control for continuous-time nonlinear systems using a novel identifier-actor-optimizer (IAO) policy learning architecture. In this IAO-based policy learning approach, a dynamical identifier is developed to approximate the unknown part of system dynamics using deep neural networks (DNNs). Then, an indirect-method-based optimizer is proposed to generate high-quality optimal actions for system control considering both the constraints and performance index. Furthermore, a DNN-based actor is developed to approximate the obtained optimal actions and return good initial guesses to the optimizer. In this way, the traditional optimal control methods and state-of-the-art DNN techniques are combined in the IAO-based optimal policy learning method. Compared to the reinforcement learning algorithms with actor-critic architectures that suffer hard reward design and low computational efficiency, the IAO-based optimal policy learning algorithm enjoys fewer user-defined parameters, higher learning speeds, and steadier convergence properties in solving complex continuous-time optimal control problems (OCPs). Simulation results of three space flight control missions are given to substantiate the effectiveness of this IAO-based policy learning strategy and to illustrate the performance of the developed DNN-based optimal control method for continuous-time OCPs.