Deterministic Policy Gradient Algorithms

Deterministic Policy Gradient Algorithms
复制标题

DOI:
--
复制
发表时间:
2014-06
期刊:
--
影响因子:
--
通讯作者:
David Silver;Guy Lever;N. Heess;T. Degris;Daan Wierstra;Martin A. Riedmiller
David Silver;Guy Lever;N. Heess;T. Degris;Daan Wierstra;Martin A. Riedmiller
中科院分区:
其他
文献类型:
--
作者:
David Silver;Guy Lever;N. Heess;T. Degris;Daan Wierstra;Martin A. Riedmiller

文献摘要

被引文献

相似文献

在本文中,我们考虑通过连续行动进行加强学习的确定性政策梯度算法。确定性政策梯度具有特别吸引人的形式:这是行动价值功能的预期梯度。这种简单的形式意味着可以比通常的随机策略梯度更有效地估计确定性政策梯度。为了确保足够的探索,我们引入了一种非政府演员批评算法,该算法从探索性行为政策中学习确定性目标政策。我们证明,确定性的策略梯度算法可以在高维操作空间中显着胜过其随机对应物。
In this paper we consider deterministic policy gradient algorithms for reinforcement learning with continuous actions. The deterministic policy gradient has a particularly appealing form: it is the expected gradient of the action-value function. This simple form means that the deterministic policy gradient can be estimated much more efficiently than the usual stochastic policy gradient. To ensure adequate exploration, we introduce an off-policy actor-critic algorithm that learns a deterministic target policy from an exploratory behaviour policy. We demonstrate that deterministic policy gradient algorithms can significantly outperform their stochastic counterparts in high-dimensional action spaces.