Real–Sim–Real Transfer for Real-World Robot Control Policy Learning with Deep Reinforcement Learning

Real–Sim–Real Transfer for Real-World Robot Control Policy Learning with Deep Reinforcement Learning
复制标题

DOI:
10.3390/app10051555
复制
发表时间:
2020-02
期刊:
影响因子:
--
通讯作者:
N. Liu;Yinghao Cai;Tao Lu;Rui Wang;Shuo Wang
N. Liu;Yinghao Cai;Tao Lu;Rui Wang;Shuo Wang
中科院分区:
--
文献类型:
--
作者:
N. Liu;Yinghao Cai;Tao Lu;Rui Wang;Shuo Wang

文献摘要

相似文献

与传统的数据驱动学习方法相比,最近开发的深度强化学习(DRL)方法可以用于训练机器人代理以获得具有吸引力性能的控制策略。然而,通过DRL学习现实世界机器人的控制策略是昂贵和繁琐的。一个有前途的替代方案是在模拟环境中训练策略,并将学习到的策略转移到真实世界的场景中。不幸的是,由于模拟和真实世界环境之间的现实差距,在模拟环境中学习的策略往往不能很好地推广到真实的世界。弥合现实差距仍然是一个具有挑战性的问题。在本文中,我们提出了一种新的real-sim-真实的(RSR)传输方法,包括一个真正的SIM的训练阶段和一个SIM到真实的推理阶段。在real-to-sim训练阶段,基于真实场景的语义信息和坐标变换,构建与任务相关的模拟环境,并在构建的模拟环境中使用DRL方法训练策略。在模拟到真实的推理阶段,学习到的策略直接应用于在没有任何真实世界数据的真实世界场景中控制机器人。在两个不同的机器人控制任务的实验结果表明,所提出的RSR方法可以训练技能的政策,具有较高的泛化性能和显着较低的训练成本。
Compared to traditional data-driven learning methods, recently developed deep reinforcement learning (DRL) approaches can be employed to train robot agents to obtain control policies with appealing performance. However, learning control policies for real-world robots through DRL is costly and cumbersome. A promising alternative is to train policies in simulated environments and transfer the learned policies to real-world scenarios. Unfortunately, due to the reality gap between simulated and real-world environments, the policies learned in simulated environments often cannot be generalized well to the real world. Bridging the reality gap is still a challenging problem. In this paper, we propose a novel real–sim–real (RSR) transfer method that includes a real-to-sim training phase and a sim-to-real inference phase. In the real-to-sim training phase, a task-relevant simulated environment is constructed based on semantic information of the real-world scenario and coordinate transformation, and then a policy is trained with the DRL method in the built simulated environment. In the sim-to-real inference phase, the learned policy is directly applied to control the robot in real-world scenarios without any real-world data. Experimental results in two different robot control tasks show that the proposed RSR method can train skill policies with high generalization performance and significantly low training costs.