Constraint-aware learning of policies by demonstration

Constraint-aware learning of policies by demonstration
复制标题

DOI:
10.1177/0278364918784354
复制
发表时间:
2018-12-01
影响因子:
9.2
通讯作者:
Vijayakumar, Sethu
Vijayakumar, Sethu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Armesto, Leopoldo;Moura, Joao;Vijayakumar, Sethu

文献摘要

被引文献

相似文献

机器人系统中的许多实际任务,如擦窗户,书写或抓取,都受到固有的约束。受约束的学习策略是一个具有挑战性的问题。在本文中,我们提出了一种约束感知学习的方法,解决了政策学习问题,使用冗余机器人执行的政策,是在零空间的约束。特别是,我们感兴趣的是在训练过程中不知道的约束条件下推广学习到的零空间策略。我们将学习约束和策略的组合问题分为两个:首先估计约束,然后使用剩余的自由度估计零空间策略。对于一个线性参数化,我们提供了一个封闭形式的解决方案的问题。我们还定义了一个度量比较的相似性估计的约束,这是有用的预处理的轨迹记录在演示。我们已经验证了我们的方法,通过学习擦拭任务,从人类演示的平面和再现它在一个未知的曲面上使用力或扭矩为基础的控制器,以实现工具对齐。我们表明,尽管训练和验证方案之间的差异,我们学习的政策,仍然提供所需的擦拭运动。
Many practical tasks in robotic systems, such as cleaning windows, writing, or grasping, are inherently constrained. Learning policies subject to constraints is a challenging problem. In this paper, we propose a method of constraint-aware learning that solves the policy learning problem using redundant robots that execute a policy that is acting in the null space of a constraint. In particular, we are interested in generalizing learned null-space policies across constraints that were not known during the training. We split the combined problem of learning constraints and policies into two: first estimating the constraint, and then estimating a null-space policy using the remaining degrees of freedom. For a linear parametrization, we provide a closed-form solution of the problem. We also define a metric for comparing the similarity of estimated constraints, which is useful to pre-process the trajectories recorded in the demonstrations. We have validated our method by learning a wiping task from human demonstration on flat surfaces and reproducing it on an unknown curved surface using a force- or torque-based controller to achieve tool alignment. We show that, despite the differences between the training and validation scenarios, we learn a policy that still provides the desired wiping motion.