Learning manipulation skills from a single demonstration

Learning manipulation skills from a single demonstration
复制标题

DOI:
10.1177/0278364917743795
复制
发表时间:
2018
期刊:
The International Journal of Robotics Research
影响因子:
--
通讯作者:
Péter Englert;Marc Toussaint
Péter Englert;Marc Toussaint
中科院分区:
其他
文献类型:
--
作者:
Péter Englert;Marc Toussaint

文献摘要

被引文献

相似文献

我们考虑这样一个场景,一个机器人被演示了一次操作技能,然后应该只使用自己的几次试验来学习复制、优化和推广相同的技能。操作技巧通常是一种高维策略。为了达到期望的采样效率,我们需要利用该问题的固有结构。通过我们的方法,我们建议将问题分解为可分析的已知目标,例如运动平滑,以及黑盒目标,例如试验成功或奖励,这取决于与环境的交互。分解允许我们利用和结合(i)约束优化方法来解决分析目标,(ii)约束贝叶斯优化来探索黑盒目标,以及(iii)逆最优控制方法来最终提取可推广的技能表示。在一个综合基准实验中对该算法进行了评价,并与最先进的学习方法进行了比较。我们还用PR2在真实机器人实验中验证了其性能。
We consider the scenario where a robot is demonstrated a manipulation skill once and should then use only a few trials on its own to learn to reproduce, optimize, and generalize that same skill. A manipulation skill is generally a high-dimensional policy. To achieve the desired sample efficiency, we need to exploit the inherent structure in this problem. With our approach, we propose to decompose the problem into analytically known objectives, such as motion smoothness, and black-box objectives, such as trial success or reward, depending on the interaction with the environment. The decomposition allows us to leverage and combine (i) constrained optimization methods to address analytic objectives, (ii) constrained Bayesian optimization to explore black-box objectives, and (iii) inverse optimal control methods to eventually extract a generalizable skill representation. The algorithm is evaluated on a synthetic benchmark experiment and compared with state-of-the-art learning methods. We also demonstrate the performance on real-robot experiments with a PR2.