Proximal Gradient Temporal Difference Learning Algorithms

Proximal Gradient Temporal Difference Learning Algorithms
复制标题

近端梯度时间差分学习算法

DOI:
--
复制
发表时间:
2016
期刊:
International Joint Conference on Artificial Intelligence
影响因子:
--
通讯作者:
Marek Petrik
Marek Petrik
中科院分区:
--
文献类型:
--
作者:
Bo Liu;Ji Liu;M. Ghavamzadeh;S. Mahadevan;Marek Petrik

文献摘要

被引文献

相似文献

在本文中,我们描述了近似梯度时间差学习,它提供了一个原则性的方法来设计和分析真正的随机梯度时间差学习算法。我们展示了梯度TD(GTD)强化学习方法是如何正式推导出来的,不是像以前尝试的那样相对于它们的原始目标函数,而是相对于原始-对偶鞍点目标函数。我们还进行了鞍点误差分析,以获得其性能的有限样本界限。以往的分析这类算法使用随机逼近技术来证明渐近收敛,并没有有限样本分析已经尝试。还提出了一种加速算法,即GTD 2-MP,它使用邻近的“镜像映射”来产生加速。我们的理论分析的结果意味着,GTD家族的算法是可比的,可能确实是首选现有的最小二乘TD方法的离线学习,由于其线性复杂性。我们提供的实验结果表明,我们的加速梯度TD方法的性能提高。
In this paper, we describe proximal gradient temporal difference learning, which provides a principled way for designing and analyzing true stochastic gradient temporal difference learning algorithms. We show how gradient TD (GTD) reinforcement learning methods can be formally derived, not with respect to their original objective functions as previously attempted, but rather with respect to primal-dual saddle-point objective functions. We also conduct a saddle-point error analysis to obtain finite-sample bounds on their performance. Previous analyses of this class of algorithms use stochastic approximation techniques to prove asymptotic convergence, and no finite-sample analysis had been attempted. An accelerated algorithm is also proposed, namely GTD2-MP, which use proximal "mirror maps" to yield acceleration. The results of our theoretical analysis imply that the GTD family of algorithms are comparable and may indeed be preferred over existing least squares TD methods for off-policy learning, due to their linear complexity. We provide experimental results showing the improved performance of our accelerated gradient TD methods.