TAX-Pose: Task-Specific Cross-Pose Estimation for Robot Manipulation

TAX-Pose: Task-Specific Cross-Pose Estimation for Robot Manipulation
复制标题

DOI:
10.48550/arxiv.2211.09325
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Chuer Pan;Brian Okorn;Harry Zhang;Ben Eisner;David Held
Chuer Pan;Brian Okorn;Harry Zhang;Ben Eisner;David Held
中科院分区:
其他
文献类型:
--
作者:
Chuer Pan;Brian Okorn;Harry Zhang;Ben Eisner;David Held

文献摘要

相似文献

我们如何赋予机器人高效操作未见过的物体以及基于演示转移相关技能的能力呢?端到端学习方法往往无法泛化到新的物体或未见过的配置。相反,我们专注于交互物体相关部分之间特定任务的位姿关系。我们推测这种关系是一种可泛化的操作任务概念,能够转移到同一类别中的新物体上;例如平底锅相对于烤箱的位姿关系,或者马克杯相对于杯架的位姿关系。我们将这种特定任务的位姿关系称为“交叉位姿”,并对这一概念给出了数学定义。我们提出了一个基于视觉的系统,该系统利用学习到的跨物体对应关系,学习针对给定操作任务估计两个物体之间的交叉位姿。然后,估计出的交叉位姿被用于引导下游的运动规划器将物体操作到期望的位姿关系(将平底锅放入烤箱,或将马克杯放在杯架上)。我们展示了我们的方法泛化到未见过物体的能力,在某些情况下,在现实世界中仅经过10次演示训练后就具备了这种能力。结果表明,我们的系统在多个任务的模拟和现实世界实验中都达到了最先进的性能。补充信息和视频可在https://sites.google.com/view/tax - pose/home找到。
How do we imbue robots with the ability to efficiently manipulate unseen objects and transfer relevant skills based on demonstrations? End-to-end learning methods often fail to generalize to novel objects or unseen configurations. Instead, we focus on the task-specific pose relationship between relevant parts of interacting objects. We conjecture that this relationship is a generalizable notion of a manipulation task that can transfer to new objects in the same category; examples include the relationship between the pose of a pan relative to an oven or the pose of a mug relative to a mug rack. We call this task-specific pose relationship"cross-pose"and provide a mathematical definition of this concept. We propose a vision-based system that learns to estimate the cross-pose between two objects for a given manipulation task using learned cross-object correspondences. The estimated cross-pose is then used to guide a downstream motion planner to manipulate the objects into the desired pose relationship (placing a pan into the oven or the mug onto the mug rack). We demonstrate our method's capability to generalize to unseen objects, in some cases after training on only 10 demonstrations in the real world. Results show that our system achieves state-of-the-art performance in both simulated and real-world experiments across a number of tasks. Supplementary information and videos can be found at https://sites.google.com/view/tax-pose/home.