Single-Stage Keypoint- Based Category-Level Object Pose Estimation from an RGB Image

Single-Stage Keypoint- Based Category-Level Object Pose Estimation from an RGB Image
复制标题

DOI:
10.1109/icra46639.2022.9812299
复制
发表时间:
2021-09
期刊:
2022 International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Yunzhi Lin;Jonathan Tremblay;Stephen Tyree;P. Vela;Stan Birchfield
Yunzhi Lin;Jonathan Tremblay;Stephen Tyree;P. Vela;Stan Birchfield
中科院分区:
其他
文献类型:
--
作者:
Yunzhi Lin;Jonathan Tremblay;Stephen Tyree;P. Vela;Stan Birchfield

文献摘要

被引文献

相似文献

关于6-DoF对象姿态估计的先前工作主要集中在实例级处理上,其中纹理CAD模型可用于被检测的每个对象。类别级6自由度姿态估计是开发在非结构化真实世界场景中运行的机器人视觉系统的重要一步。在这项工作中,我们提出了一个单阶段,基于关键点的方法,用于类别级对象姿态估计,使用单个RGB图像作为输入,在已知类别内的未知对象实例上进行操作。所提出的网络执行2D对象检测,检测2D关键点,估计6- DoF姿态,并回归相对边界长方体尺寸。这些数量是以顺序的方式估计的,利用最近的convGRU思想将信息从更容易的任务传播到更困难的任务。我们在设计选择中倾向于简单:通用长方体顶点坐标,单级网络和单目RGB输入。我们在具有挑战性的Objectron基准上进行了广泛的实验,在3D IoU指标上优于最先进的方法(比MobilePose单阶段方法高27.6%,比相关的两阶段方法高7.1%)。
Prior work on 6-DoF object pose estimation has largely focused on instance-level processing, in which a textured CAD model is available for each object being detected. Category-level 6- DoF pose estimation represents an important step toward developing robotic vision systems that operate in unstructured, real-world scenarios. In this work, we propose a single-stage, keypoint-based approach for category-level object pose estimation that operates on unknown object instances within a known category using a single RGB image as input. The proposed network performs 2D object detection, detects 2D keypoints, estimates 6- DoF pose, and regresses relative bounding cuboid dimensions. These quantities are estimated in a sequential fashion, leveraging the recent idea of convGRU for propagating information from easier tasks to those that are more difficult. We favor simplicity in our design choices: generic cuboid vertex coordinates, single-stage network, and monocular RGB input. We conduct extensive experiments on the challenging Objectron benchmark, outperforming state-of-the-art methods on the 3D IoU metric (27.6% higher than the MobilePose single-stage approach and 7.1 % higher than the related two-stage approach).