Variance-Aware Off-Policy Evaluation with Linear Function Approximation

Variance-Aware Off-Policy Evaluation with Linear Function Approximation
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yifei Min;Tianhao Wang;Dongruo Zhou;Quanquan Gu
Yifei Min;Tianhao Wang;Dongruo Zhou;Quanquan Gu
中科院分区:
其他
文献类型:
--
作者:
Yifei Min;Tianhao Wang;Dongruo Zhou;Quanquan Gu

文献摘要

相似文献

我们利用线性函数逼近研究强化学习中的离策略评估(OPE)问题,其目的是根据行为策略收集的离线数据来估计目标策略的价值函数。我们建议结合价值函数的方差信息来提高OPE的样本效率。更具体地说,对于时间不均匀的情景线性马尔可夫决策过程(MDP),我们提出了一种算法 VA-OPE,该算法使用价值函数的估计方差来重新加权拟合 Q 迭代中的贝尔曼残差。我们表明,我们的算法实现了比最著名的结果更严格的误差范围。我们还提供了行为策略和目标策略之间的分布转变的细粒度特征。大量的数值实验证实了我们的理论。
We study the off-policy evaluation (OPE) problem in reinforcement learning with linear function approximation, which aims to estimate the value function of a target policy based on the offline data collected by a behavior policy. We propose to incorporate the variance information of the value function to improve the sample efficiency of OPE. More specifically, for time-inhomogeneous episodic linear Markov decision processes (MDPs), we propose an algorithm, VA-OPE, which uses the estimated variance of the value function to reweight the Bellman residual in Fitted Q-Iteration. We show that our algorithm achieves a tighter error bound than the best-known result. We also provide a fine-grained characterization of the distribution shift between the behavior policy and the target policy. Extensive numerical experiments corroborate our theory.