Optimal Tracking Control of Unknown Discrete-Time Linear Systems Using Input-Output Measured Data

Optimal Tracking Control of Unknown Discrete-Time Linear Systems Using Input-Output Measured Data
复制标题

DOI:
10.1109/tcyb.2014.2384016
复制
发表时间:
2015-01
影响因子:
11.8
通讯作者:
Bahare Kiumarsi-Khomartash;F. Lewis;M. Naghibi-Sistani;A. Karimpour
Bahare Kiumarsi-Khomartash;F. Lewis;M. Naghibi-Sistani;A. Karimpour
中科院分区:
计算机科学1区
文献类型:
--
作者:
Bahare Kiumarsi-Khomartash;F. Lewis;M. Naghibi-Sistani;A. Karimpour

文献摘要

被引文献

相似文献

针对未知离散时间系统的无穷水平线性二次型跟踪(LQT)问题,提出了一种输出反馈解。构造了由系统动力学和参考弹道动力学组成的增广系统。扩充系统的状态由扩充系统历史中的过去输入、输出和参考轨迹的有限数量的测量来构建。建立了一个新的Bellman方程,该方程仅使用来自增广系统的输入、输出和参考轨迹数据来计算与固定策略相关的价值函数。通过使用一类强化学习方法--近似动态规划,LQT问题只需测量来自增广系统的输入、输出和参考轨迹即可在线求解,而不需要增广系统动力学知识。我们开发了策略迭代(PI)和值迭代(VI)算法,它们收敛到只需要测量输入、输出和参考轨迹数据的最优控制器。文中还给出了PI和VI算法的收敛结果。仿真算例验证了该控制方案的有效性。
In this paper, an output-feedback solution to the infinite-horizon linear quadratic tracking (LQT) problem for unknown discrete-time systems is proposed. An augmented system composed of the system dynamics and the reference trajectory dynamics is constructed. The state of the augmented system is constructed from a limited number of measurements of the past input, output, and reference trajectory in the history of the augmented system. A novel Bellman equation is developed that evaluates the value function related to a fixed policy by using only the input, output, and reference trajectory data from the augmented system. By using approximate dynamic programming, a class of reinforcement learning methods, the LQT problem is solved online without requiring knowledge of the augmented system dynamics only by measuring the input, output, and reference trajectory from the augmented system. We develop both policy iteration (PI) and value iteration (VI) algorithms that converge to an optimal controller that require only measuring the input, output, and reference trajectory data. The convergence of the proposed PI and VI algorithms is shown. A simulation example is used to verify the effectiveness of the proposed control scheme.