Online policy iteration ADP-based attitude-tracking control for hypersonic vehicles

Online policy iteration ADP-based attitude-tracking control for hypersonic vehicles
复制标题

基于ADP的在线策略迭代高超声速飞行器姿态跟踪控制

DOI:
10.1016/j.ast.2020.106233
复制
发表时间:
2020
影响因子:
5.6
通讯作者:
Yongji Wang
Yongji Wang
中科院分区:
工程技术1区
文献类型:
--
作者:
Xiao Han;Zongzhun Zheng;Lei Liu;Bo Wang;Zhongtao Cheng;Huijin Fan;Yongji Wang

文献摘要

相似文献

提出一种基于策略迭代的在线自适应动态规划(ADP)姿态跟踪控制器,旨在实现高超声速飞行器(HV)的最优控制。贝尔曼方程被称为主要的递归动态规划公式,被提供来获得控制器。特别地,控制动作由ADP控制器产生以跟踪姿态轨迹。为了在不确定的非线性高压系统中实现最优控制,我们使用策略迭代来逼近贝尔曼方程,并构建了一个行动者-预测器-批评者框架,其中采用行动网络、状态估计器和批评者网络来实现策略迭代。同时,提供离线学习方法,逼近迭代计算的初值,提高在线学习的效率。比较模拟证明了 PIADP 在气动参数扰动和随机扰动下的良好性能。
An online adaptive dynamic programming (ADP) attitude-tracking controller based on policy iteration is proposed, aiming to approach the optimal control of hypersonic vehicles (HVs). The Bellman equation, known as the principal recursive dynamic programming formula, is provided to obtain the controller. In particular, the control action is generated by the ADP controller to track the attitude trajectory. In order to approach optimal control in the uncertain nonlinear HVs system, we use policy iteration to approximate the Bellman equation and build an actor-predictor-critic framework, in which the action network, state estimator and critic network are adopted to implement the policy iteration. Meanwhile, an offline learning method is provided to approach the initial value of iterative computations and improve the efficiency of online learning. The comparative simulations demonstrate the good performance of PIADP with aerodynamic parameter perturbations and random disturbances.