Online Bootstrap Inference For Policy Evaluation In Reinforcement Learning

Online Bootstrap Inference For Policy Evaluation In Reinforcement Learning
复制标题

DOI:
10.1080/01621459.2022.2096620
复制
发表时间:
2021-08
影响因子:
3.7
通讯作者:
Pratik Ramprasad;Yuantong Li;Zhuoran Yang;Zhaoran Wang;W. Sun;Guang Cheng
Pratik Ramprasad;Yuantong Li;Zhuoran Yang;Zhaoran Wang;W. Sun;Guang Cheng
中科院分区:
数学1区
文献类型:
--
作者:
Pratik Ramprasad;Yuantong Li;Zhuoran Yang;Zhaoran Wang;W. Sun;Guang Cheng

文献摘要

被引文献

相似文献

摘要最近出现的强化学习(RL)产生了强大的统计推断方法的参数估计使用这些算法计算的需求。在线学习中的现有推理方法仅限于涉及独立采样观察的设置,而RL中的推理方法迄今为止仅限于批量设置。自举是在线学习算法中统计推断的一种灵活有效的方法,但其在涉及马尔可夫噪声(如RL)的设置中的有效性尚未得到探索。在这篇文章中,我们研究了使用在线引导方法在RL策略评估中进行推理。特别是,我们专注于时间差(TD)学习和梯度TD(GTD)学习算法,这本身就是马尔可夫噪声下的线性随机逼近的特殊情况。该方法被证明是分布一致的政策评估中的统计推断,并包括数值实验,以证明该算法在一系列真实的RL环境的有效性。本文的补充材料可在网上查阅。
Abstract The recent emergence of reinforcement learning (RL) has created a demand for robust statistical inference methods for the parameter estimates computed using these algorithms. Existing methods for inference in online learning are restricted to settings involving independently sampled observations, while inference methods in RL have so far been limited to the batch setting. The bootstrap is a flexible and efficient approach for statistical inference in online learning algorithms, but its efficacy in settings involving Markov noise, such as RL, has yet to be explored. In this article, we study the use of the online bootstrap method for inference in RL policy evaluation. In particular, we focus on the temporal difference (TD) learning and Gradient TD (GTD) learning algorithms, which are themselves special instances of linear stochastic approximation under Markov noise. The method is shown to be distributionally consistent for statistical inference in policy evaluation, and numerical experiments are included to demonstrate the effectiveness of this algorithm across a range of real RL environments. Supplementary materials for this article are available online.