Sample Efficient Policy Gradient Methods with Recursive Variance Reduction

Sample Efficient Policy Gradient Methods with Recursive Variance Reduction
复制标题

DOI:
--
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Pan Xu;F. Gao;Quanquan Gu
Pan Xu;F. Gao;Quanquan Gu
中科院分区:
其他
文献类型:
--
作者:
Pan Xu;F. Gao;Quanquan Gu

文献摘要

被引文献

相似文献

提高强化学习中的样本效率一直是一个长期存在的研究问题。在这项工作中,我们的目标是降低现有策略梯度方法的样本复杂性。我们提出了一种名为 SRVR-PG 的新颖策略梯度算法,该算法仅需要 $O(1/\epsilon^{3/2})$ 集即可找到非凹性能函数 $J(\boldsymbol{\theta})$ 的 $\epsilon$ 近似驻点(即 $\boldsymbol{\theta}$ 使得 $\|\nabla J(\boldsymbol{\theta})\|_2^2\leq\epsilon$)。此样本复杂度将随机方差减少策略梯度算法的现有结果 $O(1/\epsilon^{5/3})$ 提高了 $O(1/\epsilon^{1/6})$ 倍。此外,我们还提出了带有参数探索的 SRVR-PG 变体,它从先验概率分布中探索初始策略参数。我们对强化学习中的经典控制问题进行数值实验,以验证我们提出的算法的性能。
Improving the sample efficiency in reinforcement learning has been a long-standing research problem. In this work, we aim to reduce the sample complexity of existing policy gradient methods. We propose a novel policy gradient algorithm called SRVR-PG, which only requires $O(1/\epsilon^{3/2})$ episodes to find an $\epsilon$-approximate stationary point of the nonconcave performance function $J(\boldsymbol{\theta})$ (i.e., $\boldsymbol{\theta}$ such that $\|\nabla J(\boldsymbol{\theta})\|_2^2\leq\epsilon$). This sample complexity improves the existing result $O(1/\epsilon^{5/3})$ for stochastic variance reduced policy gradient algorithms by a factor of $O(1/\epsilon^{1/6})$. In addition, we also propose a variant of SRVR-PG with parameter exploration, which explores the initial policy parameter from a prior probability distribution. We conduct numerical experiments on classic control problems in reinforcement learning to validate the performance of our proposed algorithms.