A Generalized Path Integral Control Approach to Reinforcement Learning

A Generalized Path Integral Control Approach to Reinforcement Learning
复制标题

DOI:
10.5555/1756006.1953033
复制
发表时间:
2010-03
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Evangelos A. Theodorou;J. Buchli;S. Schaal
Evangelos A. Theodorou;J. Buchli;S. Schaal
中科院分区:
其他
文献类型:
--
作者:
Evangelos A. Theodorou;J. Buchli;S. Schaal

文献摘要

被引文献

相似文献

强化学习(RL)的目标是生成可扩展性更强、效率更高、开放参数更少的算法,它将最优控制和动态规划中的经典技术与统计估计理论中的现代学习技术相结合。在这种背景下,本文建议使用路径积分的随机最优控制框架来推导出一种新的策略参数化的随机最优控制方法。当基于随机Hamilton-Jacobi-Bellman(HJB)方程的值函数估计和最优控制是坚实的基础上,策略改进可以转化为路径积分的逼近问题,该路径积分除了探测噪声之外没有开放的算法参数。所得到的算法可以被认为是基于模型的、半基于模型的,甚至是无模型的,这取决于学习问题的结构。更新方程没有数值不稳定的危险,因为既不需要矩阵求逆,也不需要梯度学习率。我们的新算法在概率匹配的框架下与以前的RL研究显示了有趣的相似之处,并直观地解释了为什么略微启发式的概率匹配方法实际上可以执行得很好。实验评估表明,相对于基于梯度的策略学习和对高维控制问题的可伸缩性,该算法具有显著的性能改进。最后,通过一个模拟的12自由度机器狗的学习实验,说明了该算法在复杂的机器人学习场景中的功能。我们相信,基于路径积分的策略改进(PI2)为基于轨迹展开的RL提供了目前最高效、数值健壮且易于实现的算法之一。
With the goal to generate more scalable algorithms with higher efficiency and fewer open parameters, reinforcement learning (RL) has recently moved towards combining classical techniques from optimal control and dynamic programming with modern learning techniques from statistical estimation theory. In this vein, this paper suggests to use the framework of stochastic optimal control with path integrals to derive a novel approach to RL with parameterized policies. While solidly grounded in value function estimation and optimal control based on the stochastic Hamilton-Jacobi-Bellman (HJB) equations, policy improvements can be transformed into an approximation problem of a path integral which has no open algorithmic parameters other than the exploration noise. The resulting algorithm can be conceived of as model-based, semi-model-based, or even model free, depending on how the learning problem is structured. The update equations have no danger of numerical instabilities as neither matrix inversions nor gradient learning rates are required. Our new algorithm demonstrates interesting similarities with previous RL research in the framework of probability matching and provides intuition why the slightly heuristically motivated probability matching approach can actually perform well. Empirical evaluations demonstrate significant performance improvements over gradient-based policy learning and scalability to high-dimensional control problems. Finally, a learning experiment on a simulated 12 degree-of-freedom robot dog illustrates the functionality of our algorithm in a complex robot learning scenario. We believe that Policy Improvement with Path Integrals (PI2) offers currently one of the most efficient, numerically robust, and easy to implement algorithms for RL based on trajectory roll-outs.