An actor-critic method using Least Squares Temporal Difference learning

An actor-critic method using Least Squares Temporal Difference learning
复制标题

使用最小二乘时间差分学习的演员批评家方法

DOI:
--
复制
发表时间:
2009
期刊:
IEEE Conference on Decision and Control
影响因子:
--
通讯作者:
Reza Moazzez Estanjini
Reza Moazzez Estanjini
中科院分区:
--
文献类型:
--
作者:
I. Paschalidis;Keyong Li;Reza Moazzez Estanjini

文献摘要

被引文献

相似文献

在本文中,我们在行动者-批评者框架中使用最小二乘时间差异(LSTD)算法,其中行动者和批评者同时操作。也就是说,批评者不是学习固定策略的价值函数或策略梯度,而是在策略缓慢变化时在一个样本路径上进行学习。先前已针对一阶 TD 算法 TD(λ) 和 TD(1) 证明了此类过程的收敛性。然而,转换为更强大的 LSTD 并不简单,因为必须针对 LSTD 情况修改步长序列上的一些条件。我们提出了一个解决方案并证明了该过程的收敛性。此外,我们将 LSTD actor-critic 应用于仓库中智能调度叉车的应用程序。
In this paper, we use a Least Squares Temporal Difference (LSTD) algorithm in an actor-critic framework where the actor and the critic operate concurrently. That is, instead of learning the value function or policy gradient of a fixed policy, the critic carries out its learning on one sample path while the policy is slowly varying. Convergence of such a process has previously been proven for the first order TD algorithms, TD(λ) and TD(1). However, the conversion to the more powerful LSTD turns out not straightforward, because some conditions on the stepsize sequences must be modified for the LSTD case. We propose a solution and prove the convergence of the process. Furthermore, we apply the LSTD actor-critic to an application of intelligently dispatching forklifts in a warehouse.