Adaptive linear quadratic control using policy iteration

Adaptive linear quadratic control using policy iteration
复制标题

DOI:
10.1109/acc.1994.735224
复制
发表时间:
1994-06
期刊:
Proceedings of 1994 American Control Conference - ACC '94
影响因子:
--
通讯作者:
Steven J. Bradtke;B. Ydstie;A. Barto
Steven J. Bradtke;B. Ydstie;A. Barto
中科院分区:
其他
文献类型:
--
作者:
Steven J. Bradtke;B. Ydstie;A. Barto

文献摘要

被引文献

相似文献

在本文中,我们提出的稳定性和收敛性的结果,基于动态规划的强化学习应用于线性二次型调节(LQR)。我们分析的具体算法是基于Q-学习,它被证明是收敛到一个最优的控制器,提供的基础系统是可控的,一个特定的信号矢量是持续激励。这是连续问题的基于DP的强化学习算法的第一个收敛结果。
In this paper we present the stability and convergence results for dynamic programming-based reinforcement learning applied to linear quadratic regulation (LQR). The specific algorithm we analyze is based on Q-learning and it is proven to converge to an optimal controller provided that the underlying system is controllable and a particular signal vector is persistently excited. This is the first convergence result for DP-based reinforcement learning algorithms for a continuous problem.