Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function Approximation

Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function Approximation
复制标题

DOI:
--
复制
发表时间:
2019-11
期刊:
--
影响因子:
--
通讯作者:
Shangtong Zhang;Bo Liu;Hengshuai Yao;Shimon Whiteson
Shangtong Zhang;Bo Liu;Hengshuai Yao;Shimon Whiteson
中科院分区:
其他
文献类型:
--
作者:
Shangtong Zhang;Bo Liu;Hengshuai Yao;Shimon Whiteson

文献摘要

被引文献

相似文献

我们提出了第一个可证明收敛的两个时间尺度的非策略行动者-评论家算法(COF-PAC)与函数逼近。COF-PAC的关键是引入了一种新的评论家,强调评论家,它是通过梯度强调学习(GEM)训练的,梯度时间差学习和强调时间差学习的关键思想的新组合。在重点评论家和标准值函数评论家的帮助下,我们证明了COF-PAC的收敛性,其中评论家是线性的,演员可以是非线性的。
We present the first provably convergent two-timescale off-policy actor-critic algorithm (COF-PAC) with function approximation. Key to COF-PAC is the introduction of a new critic, the emphasis critic, which is trained via Gradient Emphasis Learning (GEM), a novel combination of the key ideas of Gradient Temporal Difference Learning and Emphatic Temporal Difference Learning. With the help of the emphasis critic and the canonical value function critic, we show convergence for COF-PAC, where the critics are linear and the actor can be nonlinear.