Generalized TD Learning
Generalized TD Learning
复制标题
DOI:
10.5555/1953048.2021063
复制
发表时间:
2011-02
期刊:
影响因子:
--
通讯作者:
Tsuyoshi Ueno;S. Maeda;M. Kawanabe;S. Ishii
中科院分区:
文献类型:
--
作者:
Tsuyoshi Ueno;S. Maeda;M. Kawanabe;S. Ishii
Since the invention of temporal difference (TD) learning (Sutton, 1988), many new algorithms for model-free policy evaluation have been proposed. Although they have brought much progress in practical applications of reinforcement learning (RL), there still remain fundamental problems concerning statistical properties of the value function estimation. To solve these problems, we introduce a new framework, semiparametric statistical inference, to model-free policy evaluation. This framework generalizes TD learning and its extensions, and allows us to investigate statistical properties of both of batch and online learning procedures for the value function estimation in a unified way in terms of estimating functions. Furthermore, based on this framework, we derive an optimal estimating function with the minimum asymptotic variance and propose batch and online learning algorithms which achieve the optimality.