Global Convergence of Policy Gradient Methods for the Linear Quadratic Regulator

Global Convergence of Policy Gradient Methods for the Linear Quadratic Regulator
复制标题

DOI:
--
复制
发表时间:
2018-01
影响因子:
8.7
通讯作者:
Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi
Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi

文献摘要

被引文献

相似文献

由于各种原因,用于强化学习和连续控制问题的直接政策梯度方法是一种流行的方法:1)它们在没有明确了解基础模型的情况下易于实施2)它们是一种“端到端”方法,直接优化感兴趣的性能度量标准3)他们固有地允许参数化的策略。一个值得注意的缺点是,即使在最基本的连续控制问题(线性二次调节器)中,这些方法也必须解决非凸优化问题,从计算和统计角度来看,几乎没有理解其效率。相比之下,最佳控制理论中的系统识别和基于模型的计划具有更加坚实的理论基础,在其计算和统计属性方面,人们已经知道很多。这项工作桥接了这一差距,表明(无模型)策略梯度方法全球融合到最佳解决方案,并且在其样本和计算复杂性方面有效(在相关问题依赖性数量上是多个方面的)。
Direct policy gradient methods for reinforcement learning and continuous control problems are a popular approach for a variety of reasons: 1) they are easy to implement without explicit knowledge of the underlying model 2) they are an "end-to-end" approach, directly optimizing the performance metric of interest 3) they inherently allow for richly parameterized policies. A notable drawback is that even in the most basic continuous control problem (that of linear quadratic regulators), these methods must solve a non-convex optimization problem, where little is understood about their efficiency from both computational and statistical perspectives. In contrast, system identification and model based planning in optimal control theory have a much more solid theoretical footing, where much is known with regards to their computational and statistical properties. This work bridges this gap showing that (model free) policy gradient methods globally converge to the optimal solution and are efficient (polynomially so in relevant problem dependent quantities) with regards to their sample and computational complexities.