Improper Reinforcement Learning with Gradient-based Policy Optimization

Improper Reinforcement Learning with Gradient-based Policy Optimization
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
--
影响因子:
--
通讯作者:
Mohammadi Zaki;Avinash Mohan;Aditya Gopalan;Shie Mannor
Mohammadi Zaki;Avinash Mohan;Aditya Gopalan;Shie Mannor
中科院分区:
其他
文献类型:
--
作者:
Mohammadi Zaki;Avinash Mohan;Aditya Gopalan;Shie Mannor

文献摘要

相似文献

我们考虑了一个不适当的强化学习设置,其中为未知的马尔可夫决策过程提供了一个学习器$M$基本控制器,并希望将它们最优地组合起来,以产生一个潜在的新控制器,该控制器可以优于每个基本控制器。这对于调整控制器非常有用,在不匹配或模拟环境中学习,以相对较少的试验获得给定目标环境的良好控制器。 \par 我们提出了一种基于梯度的方法,该方法在一类控制器的不适当混合上运行。我们推导出了该方法的收敛速率保证。混合物的值函数及其梯度可能不能以封闭形式提供;然而,我们表明我们可以使用滚动和同步扰动随机逼近(SPSA)来进行显式梯度下降优化。在(i)稳定倒立摆的标准控制理论基准和(ii)约束排队任务上的数值结果表明,不当的策略优化算法即使在其处理的基本策略不稳定\footnote{正在审查中。请不要散布。}时也能稳定系统。
We consider an improper reinforcement learning setting where a learner is given $M$ base controllers for an unknown Markov decision process, and wishes to combine them optimally to produce a potentially new controller that can outperform each of the base ones. This can be useful in tuning across controllers, learnt possibly in mismatched or simulated environments, to obtain a good controller for a given target environment with relatively few trials. \par We propose a gradient-based approach that operates over a class of improper mixtures of the controllers. We derive convergence rate guarantees for the approach assuming access to a gradient oracle. The value function of the mixture and its gradient may not be available in closed-form; however, we show that we can employ rollouts and simultaneous perturbation stochastic approximation (SPSA) for explicit gradient descent optimization. Numerical results on (i) the standard control theoretic benchmark of stabilizing an inverted pendulum and (ii) a constrained queueing task show that our improper policy optimization algorithm can stabilize the system even when the base policies at its disposal are unstable\footnote{Under review. Please do not distribute.}.