Reinforcement Learning for Adaptive Optimal Stationary Control of Linear Stochastic Systems

Reinforcement Learning for Adaptive Optimal Stationary Control of Linear Stochastic Systems
复制标题

DOI:
10.1109/tac.2022.3172250
复制
发表时间:
2021-07
影响因子:
6.8
通讯作者:
Bo Pang;Zhong-Ping Jiang
Bo Pang;Zhong-Ping Jiang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bo Pang;Zhong-Ping Jiang

文献摘要

被引文献

相似文献

本文研究了具有加性和乘性噪声的连续时间线性随机系统的自适应最优平稳控制问题。基于策略迭代,提出了一种新的非策略强化学习算法--基于乐观最小二乘的策略迭代算法.该算法能够从初始容许控制策略出发,直接从输入/状态数据中迭代寻找自适应最优平稳控制问题的近优策略,而无需显式地识别任何系统矩阵.在较弱的条件下,证明了所提出的基于最小二乘的策略迭代算法的解以概率1收敛到最优解的一个小邻域内。通过对三级倒立摆系统的仿真,验证了该算法的可行性和有效性。
This article studies the adaptive optimal stationary control of continuous-time linear stochastic systems with both additive and multiplicative noises, using reinforcement learning techniques. Based on policy iteration, a novel off-policy reinforcement learning algorithm, named optimistic least-squares-based policy iteration, is proposed, which is able to find iteratively near-optimal policies of the adaptive optimal stationary control problem directly from input/state data without explicitly identifying any system matrices, starting from an initial admissible control policy. The solutions given by the proposed optimistic least-squares-based policy iteration are proved to converge to a small neighborhood of the optimal solution with probability one, under mild conditions. The application of the proposed algorithm to a triple inverted pendulum example validates its feasibility and effectiveness.