Self-Supervised Interactive Object Segmentation Through a Singulation-and-Grasping Approach

Self-Supervised Interactive Object Segmentation Through a Singulation-and-Grasping Approach
复制标题

DOI:
10.48550/arxiv.2207.09314
复制
发表时间:
2022-07
影响因子:
9.5
通讯作者:
Houjian Yu;Changhyun Choi
Houjian Yu;Changhyun Choi
中科院分区:
材料科学2区
文献类型:
--
作者:
Houjian Yu;Changhyun Choi

文献摘要

相似文献

在非结构化环境中,不可见对象的实例分割是一个具有挑战性的问题。为了解决这一问题,我们提出了一种机器人学习方法,该方法主动与新对象交互,收集每个对象的训练标签进行进一步微调,以提高分割模型的性能,同时避免了人工标记数据集的耗时过程。通过端到端强化学习来训练单点和抓取(SAG)策略。对于一堆杂乱的物体,我们的方法选择推动和抓取运动来打破杂乱,并进行与物体无关的抓取,SAG策略将视觉观察和不完美的分割作为输入。我们将问题分解为三个子任务:(1)目标分离子任务旨在将目标彼此分离,这为(2)无碰撞抓取子任务创造了更大的空间;(3)掩模生成子任务通过使用基于光流的二进制分类器和运动线索后处理来获得自标记的地面真实掩模。我们的系统在模拟的杂乱场景中达到了70%的检测成功率。该系统对玩具积木、模拟YCB物体和现实世界中的新奇物体的交互分割平均准确率分别达到87.8%、73.9%和69.3%,超过了几条基线。
Instance segmentation with unseen objects is a challenging problem in unstructured environments. To solve this problem, we propose a robot learning approach to actively interact with novel objects and collect each object's training label for further fine-tuning to improve the segmentation model performance, while avoiding the time-consuming process of manually labeling a dataset. The Singulation-and-Grasping (SaG) policy is trained through end-to-end reinforcement learning. Given a cluttered pile of objects, our approach chooses pushing and grasping motions to break the clutter and conducts object-agnostic grasping for which the SaG policy takes as input the visual observations and imperfect segmentation. We decompose the problem into three subtasks: (1) the object singulation subtask aims to separate the objects from each other, which creates more space that alleviates the difficulty of (2) the collision-free grasping subtask; (3) the mask generation subtask to obtain the self-labeled ground truth masks by using an optical flow-based binary classifier and motion cue post-processing for transfer learning. Our system achieves 70% singulation success rate in simulated cluttered scenes. The interactive segmentation of our system achieves 87.8%, 73.9%, and 69.3% average precision for toy blocks, YCB objects in simulation and real-world novel objects, respectively, which outperforms several baselines.