SimGAN: Hybrid Simulator Identification for Domain Adaptation via Adversarial Reinforcement Learning

SimGAN: Hybrid Simulator Identification for Domain Adaptation via Adversarial Reinforcement Learning
复制标题

DOI:
10.1109/icra48506.2021.9561731
复制
发表时间:
2021-01
期刊:
2021 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Yifeng Jiang;Tingnan Zhang;Daniel Ho;Yunfei Bai;C. Liu;S. Levine;Jie Tan
Yifeng Jiang;Tingnan Zhang;Daniel Ho;Yunfei Bai;C. Liu;S. Levine;Jie Tan
中科院分区:
其他
文献类型:
--
作者:
Yifeng Jiang;Tingnan Zhang;Daniel Ho;Yunfei Bai;C. Liu;S. Levine;Jie Tan

文献摘要

被引文献

相似文献

随着基于学习的方法朝着机器人控制器设计自动化的方向发展,将学习到的策略转移到具有不同动态的新域(例如,从SIM到REAL的转移)仍然需要人工努力。本文介绍了SimGAN框架,该框架通过识别混合物理模拟器来匹配模拟的轨迹和来自目标领域的轨迹,使用学习的区分损失来解决与人工损失设计相关的限制,从而解决领域适应问题。我们的混合模拟器将神经网络和传统的物理模拟结合起来,以平衡表现力和泛化能力,并减少了对系统ID中精心选择的参数集的需要。一旦通过对抗性强化学习识别出混合模拟器,它就可以用于精炼目标领域的策略,而不需要交织数据收集和策略精化。我们的结果表明,我们的方法在六个领域自适应机器人运动任务上的表现优于多条强基线。
As learning-based approaches progress towards automating robot controllers design, transferring learned policies to new domains with different dynamics (e.g. sim-to-real transfer) still demands manual effort. This paper introduces SimGAN, a framework to tackle domain adaptation by identifying a hybrid physics simulator to match the simulated trajectories to the ones from the target domain, using a learned discriminative loss to address the limitations associated with manual loss design. Our hybrid simulator combines neural networks and traditional physics simulation to balance expressiveness and generalizability, and alleviates the need for a carefully selected parameter set in System ID. Once the hybrid simulator is identified via adversarial reinforcement learning, it can be used to refine policies for the target domain, without the need to interleave data collection and policy refinement. We show that our approach outperforms multiple strong baselines on six robotic locomotion tasks for domain adaptation.