Efficient Off-Policy Q-Learning for Data-Based Discrete-Time LQR Problems

Efficient Off-Policy Q-Learning for Data-Based Discrete-Time LQR Problems
复制标题

针对基于数据的离散时间 LQR 问题的高效离策略 Q 学习

DOI:
--
复制
发表时间:
2021
影响因子:
6.8
通讯作者:
M. Müller
M. Müller
中科院分区:
计算机科学2区
文献类型:
--
作者:
V. Lopez;M. Alsalti;M. Müller

文献摘要

参考文献

被引文献

相似文献

介绍并分析了一种适用于离散线性定常系统的改进Q学习算法。所提出的方法不需要任何知识的系统动力学,它享有显着的效率优于其他基于数据的最优控制方法在文献中。该算法可以完全离线执行,因为它不需要像在策略算法中那样将最优输入的当前估计应用于系统。结果表明,PE输入,定义从一个容易测试的矩阵秩条件,保证算法的收敛性。提出了一种基于数据的方法来设计算法所需的初始稳定反馈增益。在存在噪声测量的算法的鲁棒性进行了分析。我们在仿真中比较了所提出的算法不同的直接和间接的基于数据的控制设计方法。
This article introduces and analyzes an improved Q-learning algorithm for discrete-time linear time-invariant systems. The proposed method does not require any knowledge of the system dynamics, and it enjoys significant efficiency advantages over other data-based optimal control methods in the literature. This algorithm can be fully executed offline, as it does not require to apply the current estimate of the optimal input to the system as in on-policy algorithms. It is shown that a PE input, defined from an easily tested matrix rank condition, guarantees the convergence of the algorithm. A data-based method is proposed to design the initial stabilizing feedback gain that the algorithm requires. Robustness of the algorithm in the presence of noisy measurements is analyzed. We compare the proposed algorithm in simulation to different direct and indirect data-based control design methods.
DOI: --
发表时间: 2018-01
影响因子: 8.7
作者:
Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi
通讯作者: Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi