An analysis of temporal-difference learning with function approximation

An analysis of temporal-difference learning with function approximation
复制标题

DOI:
10.1109/9.580874
复制
发表时间:
1997-05-01
影响因子:
6.8
通讯作者:
VanRoy, B
VanRoy, B
中科院分区:
计算机科学2区
文献类型:
--
作者:
Tsitsiklis, JN;VanRoy, B

文献摘要

被引文献

相似文献

本文讨论了时间差分学习算法,并将其应用于无限时域折扣马尔可夫链的cost-to-go函数的逼近,分析了在有限或无限状态空间的不可约非周期马尔可夫链的单个无穷轨道上,线性函数逼近器的参数在线更新,给出了收敛性证明(概率为1),收敛极限的表征,以及由此产生的近似误差的界限,此外,我们的分析是基于一种新的推理,它提供了关于时间差异学习动态的新的直觉。除了证明新的和更强的积极结果之外,首先,我们证明了当更新不是基于马尔可夫链的轨迹时可能发生发散,这更好地调和了文献中讨论的关于时间差学习的可靠性的正面和负面结果,其次,我们给出了一个示例,该示例示出了当在存在非线性函数逼近器的情况下使用时间差学习时发散的可能性。
We discuss the temporal-difference learning algorithm, as applied to approximating the cost-to-go function of an infinite-horizon discounted Markov chain, The algorithm we analyze updates parameters of a linear function approximator online during a single endless trajectory of an irreducible aperiodic Markov chain with a finite or infinite state space, We present a proof of convergence (with probability one), a characterization of the limit of convergence, and a bound on the resulting approximation error, Furthermore, our analysis is based on a new line of reasoning that provides new intuition about the dynamics of temporal difference learning.In addition to proving new and stronger positive results than those previously available, we identify the significance of online updating and potential hazards associated with the use of nonlinear function approximators, First, we prove that divergence may occur when updates are not based on trajectories of the Markov chain, This bet reconciles positive and negative results that have been discussed in the literature, regarding the soundness of temporal-difference learning, Second, we present an example illustrating the possibility of divergence when temporal-difference learning is used in the presence of a nonlinear function approximator.