Statistically Efficient Off-Policy Policy Gradients

Statistically Efficient Off-Policy Policy Gradients
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
--
影响因子:
--
通讯作者:
Nathan Kallus;Masatoshi Uehara
Nathan Kallus;Masatoshi Uehara
中科院分区:
其他
文献类型:
--
作者:
Nathan Kallus;Masatoshi Uehara

文献摘要

被引文献

相似文献

强化学习中的策略梯度方法通过沿估计的策略值梯度的方向采取步骤来更新策略参数。在本文中,我们考虑从非策略数据中统计有效地估计策略梯度,其中估计特别非平凡。我们推导了马尔可夫决策过程和非马尔可夫决策过程中可行均方误差的渐近下界,并证明了现有的估计器在一般情况下无法实现它。我们提出了一种元算法,该算法在没有任何参数假设的情况下达到下界,并表现出独特的三向双鲁棒性。我们讨论了如何估计算法所依赖的干扰。最后,我们建立了当我们朝着新估计的政策梯度的方向采取步骤时,我们接近平稳点的速率的保证。
Policy gradient methods in reinforcement learning update policy parameters by taking steps in the direction of an estimated gradient of policy value. In this paper, we consider the statistically efficient estimation of policy gradients from off-policy data, where the estimation is particularly non-trivial. We derive the asymptotic lower bound on the feasible mean-squared error in both Markov and non-Markov decision processes and show that existing estimators fail to achieve it in general settings. We propose a meta-algorithm that achieves the lower bound without any parametric assumptions and exhibits a unique 3-way double robustness property. We discuss how to estimate nuisances that the algorithm relies on. Finally, we establish guarantees on the rate at which we approach a stationary point when we take steps in the direction of our new estimated policy gradient.