Physics-based scene-level reasoning for object pose estimation in clutter

Physics-based scene-level reasoning for object pose estimation in clutter
复制标题

DOI:
10.1177/0278364919846551
复制
发表时间:
2022-05-01
影响因子:
9.2
通讯作者:
Bekris, Kostas
Bekris, Kostas
中科院分区:
计算机科学2区
文献类型:
--
作者:
Mitash, Chaitanya;Boularias, Abdeslam;Bekris, Kostas

文献摘要

被引文献

相似文献

本文重点关注对放置在杂乱中的多个刚性物体进行基于视觉的姿态估计,特别是在涉及遮挡和物体相互搁置的情况下。鉴于深度学习的进步,最近在物体识别方面取得了进展。然而,此类工具通常需要大量的训练数据和大量的人工来标记对象。这限制了它们在机器人领域的适用性,因为解决方案必须扩展到大量物体和各种条件。此外,由于放置多个对象而产生的场景的组合性质很难在训练数据集中捕获。因此,学习的模型可能无法产生机器人操作等任务所需的所需精度水平。这项工作提出了一种姿态估计的自主过程,涵盖从数据生成到场景级推理和自学习。特别是,所提出的框架首先生成一个标记数据集,用于训练卷积神经网络(CNN)以进行杂波中的目标检测。这些检测用于指导场景级优化过程,该过程考虑杂波中存在的不同对象之间的相互作用,以输出高精度的姿态估计。此外,置信估计用于标记来自多个视图的在线真实图像,并在自学习管道中重新训练该过程。实验结果表明,这个过程能够快速地在杂乱的场景中识别出物理上一致的物体姿势,这些姿势比通过对物体的单个实例进行推理所找到的姿势更精确。此外,在自学习过程中,姿态估计的质量会随着时间的推移而提高。
This paper focuses on vision-based pose estimation for multiple rigid objects placed in clutter, especially in cases involving occlusions and objects resting on each other. Progress has been achieved recently in object recognition given advancements in deep learning. Nevertheless, such tools typically require a large amount of training data and significant manual effort to label objects. This limits their applicability in robotics, where solutions must scale to a large number of objects and variety of conditions. Moreover, the combinatorial nature of the scenes that could arise from the placement of multiple objects is difficult to capture in the training dataset. Thus, the learned models might not produce the desired level of precision required for tasks, such as robotic manipulation. This work proposes an autonomous process for pose estimation that spans from data generation to scene-level reasoning and self-learning. In particular, the proposed framework first generates a labeled dataset for training a convolutional neural network (CNN) for object detection in clutter. These detections are used to guide a scene-level optimization process, which considers the interactions between the different objects present in the clutter to output pose estimates of high precision. Furthermore, confident estimates are used to label online real images from multiple views and re-train the process in a self-learning pipeline. Experimental results indicate that this process is quickly able to identify in cluttered scenes physically consistent object poses that are more precise than those found by reasoning over individual instances of objects. Furthermore, the quality of pose estimates increases over time given the self-learning process.