What Went Wrong? Closing the Sim-to-Real Gap via Differentiable Causal Discovery

What Went Wrong? Closing the Sim-to-Real Gap via Differentiable Causal Discovery
复制标题

DOI:
10.48550/arxiv.2306.15864
复制
发表时间:
2023-06
期刊:
EPL (Europhysics Letters)
影响因子:
--
通讯作者:
Peide Huang;Xilun Zhang;Ziang Cao;Shiqi Liu;Mengdi Xu;Wenhao Ding;Jonathan M Francis;Bingqing Chen;Ding Zhao
Peide Huang;Xilun Zhang;Ziang Cao;Shiqi Liu;Mengdi Xu;Wenhao Ding;Jonathan M Francis;Bingqing Chen;Ding Zhao
中科院分区:
其他
文献类型:
--
作者:
Peide Huang;Xilun Zhang;Ziang Cao;Shiqi Liu;Mengdi Xu;Wenhao Ding;Jonathan M Francis;Bingqing Chen;Ding Zhao

文献摘要

相似文献

在模拟中训练控制策略比直接在真实机器人上训练控制策略更有吸引力,因为它允许以有效的方式探索不同的状态。然而,机器人模拟器不可避免地会表现出与现实世界的差异,从而产生不准确的结果,表现为动态模拟与现实(模拟与真实)的差距。现有文献提出通过主动修改特定模拟器参数以使模拟数据与现实世界的观察结果保持一致来缩小这一差距。然而,这组可调参数通常是手动选择的,以根据具体情况来减少搜索空间,这对于复杂系统来说很难扩展,并且需要广泛的领域知识。为了解决可扩展性问题并自动化参数调整过程,我们引入了 COMPASS,它通过发现环境参数与模拟与真实差距之间的因果关系,使模拟器与现实世界保持一致。具体来说,我们的方法学习从环境参数到模拟和现实世界机器人物体轨迹之间差异的可微映射。这种映射由同时学习的因果图控制,以帮助修剪参数的搜索空间,提供更好的可解释性,并提高对不可见参数的泛化。我们进行了实验来实现模拟到模拟和模拟到真实的转换,并表明我们的方法在几个具有挑战性的操作任务中比强基线在轨迹对齐和任务成功率方面有显着的改进。
Training control policies in simulation is more appealing than on real robots directly, as it allows for exploring diverse states in an efficient manner. Yet, robot simulators inevitably exhibit disparities from the real-world \rebut{dynamics}, yielding inaccuracies that manifest as the dynamical simulation-to-reality (sim-to-real) gap. Existing literature has proposed to close this gap by actively modifying specific simulator parameters to align the simulated data with real-world observations. However, the set of tunable parameters is usually manually selected to reduce the search space in a case-by-case manner, which is hard to scale up for complex systems and requires extensive domain knowledge. To address the scalability issue and automate the parameter-tuning process, we introduce COMPASS, which aligns the simulator with the real world by discovering the causal relationship between the environment parameters and the sim-to-real gap. Concretely, our method learns a differentiable mapping from the environment parameters to the differences between simulated and real-world robot-object trajectories. This mapping is governed by a simultaneously learned causal graph to help prune the search space of parameters, provide better interpretability, and improve generalization on unseen parameters. We perform experiments to achieve both sim-to-sim and sim-to-real transfer, and show that our method has significant improvements in trajectory alignment and task success rate over strong baselines in several challenging manipulation tasks.