End-to-end nonprehensile rearrangement with deep reinforcement learning and simulation-to-reality transfer

End-to-end nonprehensile rearrangement with deep reinforcement learning and simulation-to-reality transfer
复制标题

通过深度强化学习和模拟到现实的迁移进行端到端的非全面重排

DOI:
10.1016/j.robot.2019.06.007
复制
发表时间:
2019
期刊:
Robotics Auton. Syst.
影响因子:
--
通讯作者:
J. A. Stork
J. A. Stork
中科院分区:
--
文献类型:
--
作者:
Weihao Yuan;Kaiyu Hang;D. Kragic;M. Wang;J. A. Stork

文献摘要

参考文献

被引文献

相似文献

不可抓握重排是控制机器人通过推动作与物体进行交互,从而将物体重新配置成预定的目标姿态的问题。在这项工作中,我们使用端到端策略在有障碍物的环境中一次重新排列一个对象,该策略将原始像素映射为视觉输入来控制动作,而不需要任何形式的工程特征提取。为了减少使用真实机器人需要收集的训练数据量,我们提出了一种模拟到现实的转移方法。在第一步中,我们在仿真中对不可理解的重排任务进行建模,并使用深度强化学习来学习合适的重排策略,这需要数以十万计的示例动作进行训练。之后,我们收集了一个只有70集真实世界行为的小数据集,作为监督示例,用于使学习到的重排策略适应真实世界的输入数据。在这个过程中,我们使用了新提出的策略来改进强化学习过程,例如启发式探索和平衡经验集的管理。我们使用Baxter机器人在仿真和真实环境中对我们的方法进行了评估,结果表明,即使相机姿势与仿真不同,该方法也可以有效地改善仿真中的训练过程,并有效地使学习到的策略适应现实世界的应用。此外,我们表明,学习系统不仅可以提供自适应行为来处理执行过程中不可预见的事件,如分散注意力的物体、物体位置的突然变化和障碍物,而且还可以处理训练过程中不存在的障碍物形状。
Nonprehensile rearrangement is the problem of controlling a robot to interact with objects through pushing actions in order to reconfigure the objects into a predefined goal pose. In this work, we rearrange one object at a time in an environment with obstacles using an end-to-end policy that maps raw pixels as visual input to control actions without any form of engineered feature extraction. To reduce the amount of training data that needs to be collected using a real robot, we propose a simulation-to-reality transfer approach. In the first step, we model the nonprehensile rearrangement task in simulation and use deep reinforcement learning to learn a suitable rearrangement policy, which requires in the order of hundreds of thousands of example actions for training. Thereafter, we collect a small dataset of only 70 episodes of real-world actions as supervised examples for adapting the learned rearrangement policy to real-world input data. In this process, we make use of newly proposed strategies for improving the reinforcement learning process, such as heuristic exploration and the curation of a balanced set of experiences. We evaluate our method in both simulation and real setting using a Baxter robot to show that the proposed approach can effectively improve the training process in simulation, as well as efficiently adapt the learned policy to the real world application, even when the camera pose is different from simulation. Additionally, we show that the learned system not only can provide adaptive behavior to handle unforeseen events during executions, such as distraction objects, sudden changes in positions of the objects, and obstacles, but also can deal with obstacle shapes that were not present in the training process.
DOI: --
发表时间: 2017-07
期刊: ArXiv
影响因子: --
作者:
Stephen James;A. Davison;Edward Johns
通讯作者: Stephen James;A. Davison;Edward Johns