Fast Model Identification via Physics Engines for Data-Efficient Policy Search

Fast Model Identification via Physics Engines for Data-Efficient Policy Search
复制标题

DOI:
10.24963/ijcai.2018/451
复制
发表时间:
2017-10
期刊:
--
影响因子:
--
通讯作者:
Shaojun Zhu;A. Kimmel;Kostas E. Bekris;Abdeslam Boularias
Shaojun Zhu;A. Kimmel;Kostas E. Bekris;Abdeslam Boularias
中科院分区:
其他
文献类型:
--
作者:
Shaojun Zhu;A. Kimmel;Kostas E. Bekris;Abdeslam Boularias

文献摘要

被引文献

相似文献

本文提出了一种机器人或物体的质量、摩擦系数等力学参数的辨识方法。主要特点是使用现成的物理引擎和贝叶斯优化技术,以最大限度地减少基于模型的强化学习所需的真实世界实验的数量。所提出的框架再现在物理引擎上进行实验的真实的机器人和优化模型的机械参数,以匹配真实世界的轨迹。然后,在实际部署之前,优化的模型用于在模拟中学习策略。然而,众所周知,在仿真中很难精确地再现真实的轨迹。此外,一个接近最优的政策往往可以找到一个不完美的模型。因此,这项工作提出了一种策略,用于识别一个模型,该模型足以近似具有一定置信度的局部最优策略的值,而不是浪费精力识别最准确的模型。在仿真和真实的机器人操作任务上进行的评估表明,所提出的策略导致了一个整体的时间效率,集成的模型识别和学习解决方案,显着提高了现有的政策搜索算法的数据效率。
This paper presents a method for identifying mechanical parameters of robots or objects, such as their mass and friction coefficients. Key features are the use of off-the-shelf physics engines and the adaptation of a Bayesian optimization technique towards minimizing the number of real-world experiments needed for model-based reinforcement learning. The proposed framework reproduces in a physics engine experiments performed on a real robot and optimizes the model's mechanical parameters so as to match real-world trajectories. The optimized model is then used for learning a policy in simulation, before real-world deployment. It is well understood, however, that it is hard to exactly reproduce real trajectories in simulation. Moreover, a near-optimal policy can be frequently found with an imperfect model. Therefore, this work proposes a strategy for identifying a model that is just good enough to approximate the value of a locally optimal policy with a certain confidence, instead of wasting effort on identifying the most accurate model. Evaluations, performed both in simulation and on a real robotic manipulation task, indicate that the proposed strategy results in an overall time-efficient, integrated model identification and learning solution, which significantly improves the data-efficiency of existing policy search algorithms.