Generalized TD Learning

Generalized TD Learning
复制标题

DOI:
10.5555/1953048.2021063
复制
发表时间:
2011-02
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Tsuyoshi Ueno;S. Maeda;M. Kawanabe;S. Ishii
Tsuyoshi Ueno;S. Maeda;M. Kawanabe;S. Ishii
中科院分区:
其他
文献类型:
--
作者:
Tsuyoshi Ueno;S. Maeda;M. Kawanabe;S. Ishii

文献摘要

相似文献

自时间差分(TD)学习的发明(Sutton, 1988)以来,许多新的无模型策略评估算法被提出。虽然他们在强化学习的实际应用中取得了很大的进展,但在值函数估计的统计性质方面仍然存在一些根本性的问题。为了解决这些问题,我们将半参数统计推理引入到无模型政策评估中。该框架概括了TD学习及其扩展,并允许我们以估计函数的统一方式研究批处理和在线学习过程的值函数估计的统计性质。在此基础上,推导出具有最小渐近方差的最优估计函数,并提出了实现最优性的批处理和在线学习算法。
Since the invention of temporal difference (TD) learning (Sutton, 1988), many new algorithms for model-free policy evaluation have been proposed. Although they have brought much progress in practical applications of reinforcement learning (RL), there still remain fundamental problems concerning statistical properties of the value function estimation. To solve these problems, we introduce a new framework, semiparametric statistical inference, to model-free policy evaluation. This framework generalizes TD learning and its extensions, and allows us to investigate statistical properties of both of batch and online learning procedures for the value function estimation in a unified way in terms of estimating functions. Furthermore, based on this framework, we derive an optimal estimating function with the minimum asymptotic variance and propose batch and online learning algorithms which achieve the optimality.