A semiparametric statistical approach to model-free policy evaluation

A semiparametric statistical approach to model-free policy evaluation
复制标题

DOI:
10.1145/1390156.1390291
复制
发表时间:
2008-07
期刊:
--
影响因子:
--
通讯作者:
Tsuyoshi Ueno;M. Kawanabe;Takeshi Mori;S. Maeda;S. Ishii
Tsuyoshi Ueno;M. Kawanabe;Takeshi Mori;S. Maeda;S. Ishii
中科院分区:
其他
文献类型:
--
作者:
Tsuyoshi Ueno;M. Kawanabe;Takeshi Mori;S. Maeda;S. Ishii

文献摘要

相似文献

基于最小二乘时间差分(LSTD)的强化学习(RL)方法近年来得到了发展,并显示出良好的实用性能。然而,其估计的质量尚未得到很好的说明。本文从半参数统计推断的新视角讨论了基于LSTD的政策评估。事实上,估计量可以从特定的估计函数中获得,该函数保证其渐进收敛到真值,而无需指定环境模型。基于这些观察,我们1)分析的渐近方差的LSTD为基础的估计,2)推导出最佳的估计函数与最小的渐近估计方差,和3)推导出一个次优估计,以减少计算负担,在获得最佳的估计函数。
Reinforcement learning (RL) methods based on least-squares temporal difference (LSTD) have been developed recently and have shown good practical performance. However, the quality of their estimation has not been well elucidated. In this article, we discuss LSTD-based policy evaluation from the new view-point of semiparametric statistical inference. In fact, the estimator can be obtained from a particular estimating function which guarantees its convergence to the true value asymptotically, without specifying a model of the environment. Based on these observations, we 1) analyze the asymptotic variance of an LSTD-based estimator, 2) derive the optimal estimating function with the minimum asymptotic estimation variance, and 3) derive a suboptimal estimator to reduce the computational burden in obtaining the optimal estimating function.