Policy Gradient Method For Robust Reinforcement Learning

Policy Gradient Method For Robust Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2205.07344
复制
发表时间:
2022-05
期刊:
--
影响因子:
--
通讯作者:
Yue Wang;Shaofeng Zou
Yue Wang;Shaofeng Zou
中科院分区:
其他
文献类型:
--
作者:
Yue Wang;Shaofeng Zou

文献摘要

相似文献

本文针对模型失配情况下的稳健强化学习,提出了首个具备全局最优性保证和复杂度分析的策略梯度方法。稳健强化学习旨在学习一种对模拟器与真实环境之间的模型失配具有稳健性的策略。我们首先推导出稳健策略(次)梯度,该梯度适用于任何可微的参数化策略类别。我们证明,所提出的稳健策略梯度方法在直接策略参数化下渐近收敛到全局最优解。我们进一步开发了一种平滑的稳健策略梯度方法,并表明要达到 $\epsilon$ -全局最优解,其复杂度为 $\mathcal O(\epsilon^{-3})$ 。随后,我们将该方法扩展到一般的无模型设定,并设计了针对可微参数化策略类别和价值函数的稳健演员 - 评论家方法。我们还刻画了其在表格设定下的渐近收敛性和样本复杂度。最后,我们通过仿真结果展示了所提方法的稳健性。
This paper develops the first policy gradient method with global optimality guarantee and complexity analysis for robust reinforcement learning under model mismatch. Robust reinforcement learning is to learn a policy robust to model mismatch between simulator and real environment. We first develop the robust policy (sub-)gradient, which is applicable for any differentiable parametric policy class. We show that the proposed robust policy gradient method converges to the global optimum asymptotically under direct policy parameterization. We further develop a smoothed robust policy gradient method and show that to achieve an $\epsilon$-global optimum, the complexity is $\mathcal O(\epsilon^{-3})$. We then extend our methodology to the general model-free setting and design the robust actor-critic method with differentiable parametric policy class and value function. We further characterize its asymptotic convergence and sample complexity under the tabular setting. Finally, we provide simulation results to demonstrate the robustness of our methods.