Continuous Deep Q-Learning with Model-based Acceleration

Continuous Deep Q-Learning with Model-based Acceleration
复制标题

DOI:
--
复制
发表时间:
2016-03
期刊:
--
影响因子:
--
通讯作者:
S. Gu;T. Lillicrap;I. Sutskever;S. Levine
S. Gu;T. Lillicrap;I. Sutskever;S. Levine
中科院分区:
其他
文献类型:
--
作者:
S. Gu;T. Lillicrap;I. Sutskever;S. Levine

文献摘要

被引文献

相似文献

无模型强化学习已成功应用于一系列具有挑战性的问题,最近已扩展到处理大型神经网络策略和值函数。然而,无模型算法的样本复杂性,特别是当使用高维函数逼近器时,往往限制了它们对物理系统的适用性。在本文中,我们探索了用于降低连续控制任务的深度强化学习的样本复杂度的算法和表示。我们提出了两个互补的技术,以提高这些算法的效率。首先,我们推导出Q学习算法的一个连续变体,我们称之为归一化优势函数(NAF),作为更常用的策略梯度和演员批评方法的替代方案。NAF表示允许我们将Q学习与经验重放应用于连续任务,并大大提高了一组模拟机器人控制任务的性能。为了进一步提高我们方法的效率,我们探索了使用学习模型来加速无模型强化学习。我们表明,迭代改装的局部线性模型对此特别有效,并在这些模型适用的领域中表现出更快的学习速度。
Model-free reinforcement learning has been successfully applied to a range of challenging problems, and has recently been extended to handle large neural network policies and value functions. However, the sample complexity of modelfree algorithms, particularly when using high-dimensional function approximators, tends to limit their applicability to physical systems. In this paper, we explore algorithms and representations to reduce the sample complexity of deep reinforcement learning for continuous control tasks. We propose two complementary techniques for improving the efficiency of such algorithms. First, we derive a continuous variant of the Q-learning algorithm, which we call normalized advantage functions (NAF), as an alternative to the more commonly used policy gradient and actor-critic methods. NAF representation allows us to apply Q-learning with experience replay to continuous tasks, and substantially improves performance on a set of simulated robotic control tasks. To further improve the efficiency of our approach, we explore the use of learned models for accelerating model-free reinforcement learning. We show that iteratively refitted local linear models are especially effective for this, and demonstrate substantially faster learning on domains where such models are applicable.