Learning the model-free linear quadratic regulator via random search

Learning the model-free linear quadratic regulator via random search
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
--
影响因子:
--
通讯作者:
Hesameddin Mohammadi;M. Jovanović;M. Soltanolkotabi
Hesameddin Mohammadi;M. Jovanović;M. Soltanolkotabi
中科院分区:
其他
文献类型:
--
作者:
Hesameddin Mohammadi;M. Jovanović;M. Soltanolkotabi

文献摘要

被引文献

相似文献

无模型强化学习技术试图通过直接搜索控制器的参数空间来为未知的动态系统找到最优控制动作。这些方法的收敛行为和统计特性往往知之甚少,因为基本的优化问题的非凸性质,以及缺乏精确的梯度计算。在本文中,我们研究了具有未知状态空间参数的连续时间系统的标准有限时域线性二次调节器问题。我们提供了理论界的收敛速度和样本的随机搜索方法的复杂性。我们的研究结果表明,所需的模拟时间达到(cid:15)的准确性,在无模型设置和函数评估的总数都是O(log(1 /(cid:15)。
Model-free reinforcement learning techniques attempt to find an optimal control action for an unknown dynamical system by directly searching over the parameter space of controllers. The convergence behavior and statistical properties of these approaches are often poorly understood because of the nonconvex nature of the underlying optimization problems as well as the lack of exact gradient computation. In this paper, we examine the standard infinite-horizon linear quadratic regulator problem for continuous-time systems with unknown state-space parameters. We provide theoretical bounds on the convergence rate and sample complexity of a random search method. Our results demonstrate that the required simulation time for achieving (cid:15) -accuracy in a model-free setup and the total number of function evaluations are both of O (log (1 /(cid:15) )) .