LocaliseBot: Multi-view 3D Object Localisation with Differentiable Rendering for Robot Grasping

LocaliseBot: Multi-view 3D Object Localisation with Differentiable Rendering for Robot Grasping
复制标题

DOI:
10.1007/978-3-031-25075-0_47
复制
发表时间:
2023-11
期刊:
--
影响因子:
--
通讯作者:
Sujal Vijayaraghavan;Redwan Alqasemi;R. Dubey;Sudeep Sarkar
Sujal Vijayaraghavan;Redwan Alqasemi;R. Dubey;Sudeep Sarkar
中科院分区:
其他
文献类型:
--
作者:
Sujal Vijayaraghavan;Redwan Alqasemi;R. Dubey;Sudeep Sarkar

文献摘要

相似文献

机器人抓取通常遵循五个阶段:物体检测,物体定位,物体姿态估计,抓取姿态估计和抓取规划。我们专注于物体姿态估计。我们的方法依赖于三个信息:对象的多个视图,在这些视点的相机的外部参数,和对象的3D CAD模型。第一步涉及标准深度学习主干(FCN ResNet),以估计对象标签、语义分割和相对于相机的对象姿态的粗略估计。我们的新奇之处在于使用了一个细化模块,该模块从粗略的姿态估计开始,并通过可微分渲染进行优化。这是一种纯粹基于视觉的方法,无需点云或深度图像等其他信息。我们在ShapeNet数据集上评估了我们的物体姿态估计方法,并显示了对最先进技术的改进。我们还表明,在物体杂波室内数据集(OCID)Grasp数据集上,估计的物体姿态导致99.65%的抓取准确率,使用标准实践计算。
Robot grasp typically follows five stages: object detection, object localisation, object pose estimation, grasp pose estimation, and grasp planning. We focus on object pose estimation. Our approach relies on three pieces of information: multiple views of the object, the camera’s extrinsic parameters at those viewpoints, and 3D CAD models of objects. The first step involves a standard deep learning backbone (FCN ResNet) to estimate the object label, semantic segmentation, and a coarse estimate of the object pose with respect to the camera. Our novelty is using a refinement module that starts from the coarse pose estimate and refines it by optimisation through differentiable rendering. This is a purely vision-based approach that avoids the need for other information such as point cloud or depth images. We evaluate our object pose estimation approach on the ShapeNet dataset and show improvements over the state of the art. We also show that the estimated object pose results in 99.65% grasp accuracy with the ground truth grasp candidates on the Object Clutter Indoor Dataset (OCID) Grasp dataset, as computed using standard practice.