Learning the model-free linear quadratic regulator via random search
Learning the model-free linear quadratic regulator via random search
复制标题
DOI:
--
复制
发表时间:
2020-06
期刊:
影响因子:
--
通讯作者:
Hesameddin Mohammadi;M. Jovanović;M. Soltanolkotabi
中科院分区:
文献类型:
--
作者:
Hesameddin Mohammadi;M. Jovanović;M. Soltanolkotabi
Model-free reinforcement learning techniques attempt to find an optimal control action for an unknown dynamical system by directly searching over the parameter space of controllers. The convergence behavior and statistical properties of these approaches are often poorly understood because of the nonconvex nature of the underlying optimization problems as well as the lack of exact gradient computation. In this paper, we examine the standard infinite-horizon linear quadratic regulator problem for continuous-time systems with unknown state-space parameters. We provide theoretical bounds on the convergence rate and sample complexity of a random search method. Our results demonstrate that the required simulation time for achieving (cid:15) -accuracy in a model-free setup and the total number of function evaluations are both of O (log (1 /(cid:15) )) .