PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning Method

PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning Method
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Ziwei Guan;Tengyu Xu;Yingbin Liang
Ziwei Guan;Tengyu Xu;Yingbin Liang
中科院分区:
其他
文献类型:
--
作者:
Ziwei Guan;Tengyu Xu;Yingbin Liang

文献摘要

相似文献

强调时间差异(ETD)学习(Sutton et al., 2016)是一种利用函数逼近进行离策略价值函数评估的成功方法。尽管 ETD 已被证明可以渐近收敛到理想的值函数,但众所周知,ETD 经常遇到较大的方差,因此其样本复杂度会随着迭代次数的增加呈指数级快速增加。在这项工作中,我们提出了一种新的 ETD 方法,称为 PER-ETD(即定期重新启动-ETD),该方法仅在评估参数的每次迭代的有限周期内重新启动和更新后续跟踪。此外,PER-ETD采用了重启周期随迭代次数对数增加的设计,保证了方差和偏差之间的最佳权衡,并保持两者亚线性消失。我们证明 PER-ETD 收敛到与 ETD 相同的理想不动点,但将 ETD 的指数样本复杂度提高为多项式。我们的实验验证了 PER-ETD 的优越性能及其相对于 ETD 的优势。
Emphatic temporal difference (ETD) learning (Sutton et al., 2016) is a successful method to conduct the off-policy value function evaluation with function approximation. Although ETD has been shown to converge asymptotically to a desirable value function, it is well-known that ETD often encounters a large variance so that its sample complexity can increase exponentially fast with the number of iterations. In this work, we propose a new ETD method, called PER-ETD (i.e., PEriodically Restarted-ETD), which restarts and updates the follow-on trace only for a finite period for each iteration of the evaluation parameter. Further, PER-ETD features a design of the logarithmical increase of the restart period with the number of iterations, which guarantees the best trade-off between the variance and bias and keeps both vanishing sublinearly. We show that PER-ETD converges to the same desirable fixed point as ETD, but improves the exponential sample complexity of ETD to be polynomials. Our experiments validate the superior performance of PER-ETD and its advantage over ETD.