Overcoming Exploration in Reinforcement Learning with Demonstrations

Overcoming Exploration in Reinforcement Learning with Demonstrations
复制标题

DOI:
10.1109/icra.2018.8463162
复制
发表时间:
2017-09
期刊:
2018 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Ashvin Nair;Bob McGrew;Marcin Andrychowicz;Wojciech Zaremba;P. Abbeel
Ashvin Nair;Bob McGrew;Marcin Andrychowicz;Wojciech Zaremba;P. Abbeel
中科院分区:
其他
文献类型:
--
作者:
Ashvin Nair;Bob McGrew;Marcin Andrychowicz;Wojciech Zaremba;P. Abbeel

文献摘要

被引文献

相似文献

在奖励稀疏的环境中进行探索一直是强化学习(RL)中一个长期存在的问题。许多任务都很自然地使用稀疏奖励来指定,而手动塑造奖励函数可能会导致次优性能。然而,随着任务视界或行动维度的增加,找到非零奖励的难度呈指数级增加。这使得许多现实世界的任务超出了RL方法的实际范围。在这项工作中,我们使用演示来克服探索问题,并成功地学会了执行长视界、多步骤的机器人任务,例如用机器人手臂堆叠块。我们的方法建立在深度确定性策略梯度和事后经验回放的基础上,在模拟机器人任务上提供了比强化学习更快的速度。它很容易实现,并且只做了一个额外的假设,即我们可以收集一小组演示。此外,我们的方法能够解决无法通过RL或行为克隆单独解决的任务,并且通常最终优于演示者策略。
Exploration in environments with sparse rewards has been a persistent problem in reinforcement learning (RL). Many tasks are natural to specify with a sparse reward, and manually shaping a reward function can result in suboptimal performance. However, finding a non-zero reward is exponentially more difficult with increasing task horizon or action dimensionality. This puts many real-world tasks out of practical reach of RL methods. In this work, we use demonstrations to overcome the exploration problem and successfully learn to perform long-horizon, multi-step robotics tasks with continuous control such as stacking blocks with a robot arm. Our method, which builds on top of Deep Deterministic Policy Gradients and Hindsight Experience Replay, provides an order of magnitude of speedup over RL on simulated robotics tasks. It is simple to implement and makes only the additional assumption that we can collect a small set of demonstrations. Furthermore, our method is able to solve tasks not solvable by either RL or behavior cloning alone, and often ends up outperforming the demonstrator policy.