Optimal adaptive control for unknown systems using output feedback by reinforcement learning methods

Optimal adaptive control for unknown systems using output feedback by reinforcement learning methods
复制标题

DOI:
10.1109/icca.2010.5524211
复制
发表时间:
2010-06
期刊:
IEEE ICCA 2010
影响因子:
--
通讯作者:
F. Lewis;K. Vamvoudakis
F. Lewis;K. Vamvoudakis
中科院分区:
其他
文献类型:
--
作者:
F. Lewis;K. Vamvoudakis

文献摘要

被引文献

相似文献

最优反馈控制器通常离线计算,假设系统动态的全部知识。自适应控制器,另一方面,是在线计划,有效地学习补偿未知的系统动态和干扰。一般来说,直接自适应计划不收敛到最优控制解决方案,用户规定的性能指标。在过去的几年中,已经表明,从计算智能的强化学习技术可以用来学习最优反馈控制器在线使用直接自适应控制技术,而不知道系统的动态。大多数强化学习方法需要对系统内部状态进行全面测量。在本文中,我们开发的强化学习方法,只需要输出反馈,但收敛到最优控制器。确定性线性定常系统被认为是。策略迭代(PI)和价值迭代(VI)算法。这对应于一类部分可观测马尔可夫决策过程(POMDPs)的最优控制。结果表明,类似于Q-学习,新的输出反馈最优学习方法具有重要的优点,即不需要系统动力学的知识来实现。只需要知道系统的阶数和它的“可观测性指数”的上界。学习输出反馈控制器是多项式阿尔马控制器的形式,具有与最优状态变量反馈增益等效的性能。
Optimal feedback controllers are generally computed offline assuming full knowledge of the system dynamics. Adaptive controllers, on the other hand, are online schemes that effectively learn to compensate for unknown system dynamics and disturbances. Generally, direct adaptive schemes do not converge to optimal control solutions for user-prescribed performance measures. During the past years, it has been shown that reinforcement learning techniques from computational intelligence can be used to learn optimal feedback controllers online using direct adaptive control techniques without knowing the system dynamics. Most reinforcement learning methods require full measurements of the system internal state. In this paper we develop reinforcement learning methods which require only output feedback and yet converge to an optimal controller. Deterministic linear time-invariant systems are considered. Both policy iteration (PI) and value iteration (VI) algorithms are derived. This corresponds to optimal control for a class of partially observable Markov decision processes (POMDPs). It is shown that, similar to Q-learning, the new output-feedback optimal learning methods have the important advantage that knowledge of the system dynamics is not needed for their implementation. Only the order of the system must be known and an upper bound on its ‘observability index’. The learned output feedback controller is in the form of a polynomial ARMA controller that has equivalent performance with the optimal state variable feedback gain.