Deterministic policy gradient: Convergence analysis

Deterministic policy gradient: Convergence analysis
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Huaqing Xiong;Tengyu Xu;Lin Zhao;Yingbin Liang;Wei Zhang
Huaqing Xiong;Tengyu Xu;Lin Zhao;Yingbin Liang;Wei Zhang
中科院分区:
其他
文献类型:
--
作者:
Huaqing Xiong;Tengyu Xu;Lin Zhao;Yingbin Liang;Wei Zhang

文献摘要

相似文献

银等人[2014]提出的确定性策略梯度(DPG)方法已被证明具有上级性能,特别是对于具有多维和连续动作空间的应用。然而,目前还不清楚DPG是否收敛,如果是,它收敛的速度有多快,以及它是否像其他PG方法一样有效。在本文中,我们提供了一个理论分析的DPG回答这些问题。我们研究了单时间尺度DPG(通常是在实践中的情况下),在这两个政策和政策的设置,并表明,这两种算法达到了一个精确的静态策略,直到一个系统错误的样本复杂度为O(N-2)。此外,我们建立了高斯噪声下的DPG的收敛速度,这是广泛采用在实践中,以提高DPG的性能。据我们所知,这是DPG方法的第一个非渐近收敛特征。
The deterministic policy gradient (DPG) method proposed in Silver et al. [2014] has been demonstrated to exhibit superior performance particularly for applications with multi-dimensional and continuous action spaces. However, it remains unclear whether DPG converges, and if so, how fast it converges and whether it converges as efficiently as other PG methods. In this paper, we provide a theoretical analysis of DPG to answer those questions. We study the single timescale DPG (often the case in practice) in both on-policy and off-policy settings, and show that both algorithms attain an ϵ - accurate stationary policy up to a system error with a sample complexity of O ( ϵ − 2 ) . Moreover, we establish the convergence rate for DPG under Gaussian noise exploration, which is widely adopted in practice to improve the performance of DPG. To our best knowledge, this is the first non-asymptotic convergence characterization for DPG methods.