On actor-critic algorithms

On actor-critic algorithms
复制标题

DOI:
10.1137/s0363012901385691
复制
发表时间:
2003-01-01
影响因子:
2.2
通讯作者:
Tsitsiklis, JN
Tsitsiklis, JN
中科院分区:
数学2区
文献类型:
--
作者:
Konda, VR;Tsitsiklis, JN

文献摘要

被引文献

相似文献

在这篇文章中,我们提出并分析了一类演员评论家算法。这些都是两个时间尺度的算法,其中的评论家使用时间差学习与线性参数化的近似架构,和演员更新的近似梯度方向,根据评论家提供的信息。我们表明,批评家的功能应该理想地跨越一个子空间所规定的参数化的演员的选择。我们研究波兰状态和动作空间的马尔可夫决策过程的Actor-Critic算法。我们陈述并证明了两个关于其收敛性的结果。
In this article, we propose and analyze a class of actor-critic algorithms. These are two-time-scale algorithms in which the critic uses temporal difference learning with a linearly parameterized approximation architecture, and the actor is updated in an approximate gradient direction, based on information provided by the critic. We show that the features for the critic should ideally span a subspace prescribed by the choice of parameterization of the actor. We study actor-critic algorithms for Markov decision processes with Polish state and action spaces. We state and prove two results regarding their convergence.