Natural Policy Gradient Methods with Parameter-based Exploration for Control Tasks

Natural Policy Gradient Methods with Parameter-based Exploration for Control Tasks
复制标题

DOI:
--
复制
发表时间:
2010-12
期刊:
--
影响因子:
--
通讯作者:
Atsushi Miyamae;Y. Nagata;I. Ono;S. Kobayashi
Atsushi Miyamae;Y. Nagata;I. Ono;S. Kobayashi
中科院分区:
其他
文献类型:
--
作者:
Atsushi Miyamae;Y. Nagata;I. Ono;S. Kobayashi

文献摘要

被引文献

相似文献

在本文中,我们提出了一个有效的算法,估计自然的政策梯度,使用基于参数的探索,该算法直接在参数空间中的样本。与以前的方法基于自然梯度,我们的算法计算自然政策梯度使用的确切Fisher信息矩阵的逆。该算法的计算成本等于传统的策略梯度,而以前的自然策略梯度方法有一个令人望而却步的计算成本。实验结果表明,该方法优于几种策略梯度方法。
In this paper, we propose an efficient algorithm for estimating the natural policy gradient using parameter-based exploration; this algorithm samples directly in the parameter space. Unlike previous methods based on natural gradients, our algorithm calculates the natural policy gradient using the inverse of the exact Fisher information matrix. The computational cost of this algorithm is equal to that of conventional policy gradients whereas previous natural policy gradient methods have a prohibitive computational cost. Experimental results show that the proposed method outperforms several policy gradient methods.