Data-efficient Deep Reinforcement Learning for Dexterous Manipulation

Data-efficient Deep Reinforcement Learning for Dexterous Manipulation
复制标题

DOI:
--
复制
发表时间:
2017-04
期刊:
ArXiv
影响因子:
--
通讯作者:
I. Popov;N. Heess;T. Lillicrap;Roland Hafner;Gabriel Barth-Maron;Matej Vecerík;Thomas Lampe;Yuval Tassa;Tom Erez;Martin A. Riedmiller
I. Popov;N. Heess;T. Lillicrap;Roland Hafner;Gabriel Barth-Maron;Matej Vecerík;Thomas Lampe;Yuval Tassa;Tom Erez;Martin A. Riedmiller
中科院分区:
其他
文献类型:
--
作者:
I. Popov;N. Heess;T. Lillicrap;Roland Hafner;Gabriel Barth-Maron;Matej Vecerík;Thomas Lampe;Yuval Tassa;Tom Erez;Martin A. Riedmiller

文献摘要

被引文献

相似文献

深度学习和强化学习方法最近已被用于解决连续控制领域的各种问题。这些技术的一个明显的应用是机器人中的灵巧操作任务,这些任务很难使用传统的控制理论或手工工程方法来解决。这种任务的一个例子是抓住一个物体并将其精确地堆叠在另一个物体上。在真实的世界中解决这一困难且实际相关的问题是机器人领域的一个重要的长期目标。在这里,我们通过研究模拟中的问题并提供旨在解决它的模型和技术,向这个目标迈出了一步。我们介绍了深度确定性策略梯度算法(DDPG)的两个扩展,这是一种无模型的基于Q学习的方法,使其具有更高的数据效率和可扩展性。我们的研究结果表明,通过广泛使用的政策外的数据和重放,它是可能的控制策略,鲁棒地抓住对象和堆栈。此外,我们的研究结果暗示,通过收集真实的机器人上的交互来训练成功的堆叠策略可能很快就可行了。
Deep learning and reinforcement learning methods have recently been used to solve a variety of problems in continuous control domains. An obvious application of these techniques is dexterous manipulation tasks in robotics which are difficult to solve using traditional control theory or hand-engineered approaches. One example of such a task is to grasp an object and precisely stack it on another. Solving this difficult and practically relevant problem in the real world is an important long-term goal for the field of robotics. Here we take a step towards this goal by examining the problem in simulation and providing models and techniques aimed at solving it. We introduce two extensions to the Deep Deterministic Policy Gradient algorithm (DDPG), a model-free Q-learning based method, which make it significantly more data-efficient and scalable. Our results show that by making extensive use of off-policy data and replay, it is possible to find control policies that robustly grasp objects and stack them. Further, our results hint that it may soon be feasible to train successful stacking policies by collecting interactions on real robots.