Natural Policy Gradient Methods with Parameter-based Exploration for Control Tasks
Natural Policy Gradient Methods with Parameter-based Exploration for Control Tasks
复制标题
DOI:
--
复制
发表时间:
2010-12
期刊:
影响因子:
--
通讯作者:
Atsushi Miyamae;Y. Nagata;I. Ono;S. Kobayashi
中科院分区:
文献类型:
--
作者:
Atsushi Miyamae;Y. Nagata;I. Ono;S. Kobayashi
In this paper, we propose an efficient algorithm for estimating the natural policy gradient using parameter-based exploration; this algorithm samples directly in the parameter space. Unlike previous methods based on natural gradients, our algorithm calculates the natural policy gradient using the inverse of the exact Fisher information matrix. The computational cost of this algorithm is equal to that of conventional policy gradients whereas previous natural policy gradient methods have a prohibitive computational cost. Experimental results show that the proposed method outperforms several policy gradient methods.