Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation

Normalized Object Coordinate Space for Category-Level 6D Object Pose and Size Estimation
复制标题

DOI:
10.1109/cvpr.2019.00275
复制
发表时间:
2019-01
期刊:
2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
He Wang;Srinath Sridhar;Jingwei Huang;Julien P. C. Valentin;Shuran Song;L. Guibas
He Wang;Srinath Sridhar;Jingwei Huang;Julien P. C. Valentin;Shuran Song;L. Guibas
中科院分区:
其他
文献类型:
--
作者:
He Wang;Srinath Sridhar;Jingwei Huang;Julien P. C. Valentin;Shuran Song;L. Guibas

文献摘要

被引文献

相似文献

本文的目的是估计RGB-D图像中不可见对象实例的6D姿态和尺寸。与“实例级”的6D姿势估计任务相反,我们的问题假设在训练或测试时间内没有确切的对象CAD模型可用。为了处理给定类别中的不同和不可见的对象实例,我们引入了一个\extbf{标准化对象坐标空间(NOCS)}-一个类别中所有可能的对象实例的共享规范表示。然后训练我们的基于区域的神经网络,以直接推断从观察到的像素到该共享对象表示(NOCS)的对应关系以及其他对象信息,如类别标签和实例掩码。这些预测可以与深度图相结合,以联合估计杂乱场景中多个对象的度量6D姿势和维度。为了训练我们的网络,我们提出了一种新的上下文感知技术来生成大量完全注释的混合现实数据。为了进一步改进我们的模型并评估其在真实数据上的性能,我们还提供了一个带有大量环境和实例变化的完全标注的真实数据集。大量的实验表明,该方法能够稳健地估计真实环境中不可见对象实例的姿态和大小,同时在标准的6D姿态估计基准上也获得了最先进的性能。
The goal of this paper is to estimate the 6D pose and dimensions of unseen object instances in an RGB-D image. Contrary to ``instance-level'' 6D pose estimation tasks, our problem assumes that no exact object CAD models are available during either training or testing time. To handle different and unseen object instances in a given category, we introduce a \textbf{Normalized Object Coordinate Space (NOCS)}---a shared canonical representation for all possible object instances within a category. Our region-based neural network is then trained to directly infer the correspondence from observed pixels to this shared object representation (NOCS) along with other object information such as class label and instance mask. These predictions can be combined with the depth map to jointly estimate the metric 6D pose and dimensions of multiple objects in a cluttered scene. To train our network, we present a new context-aware technique to generate large amounts of fully annotated mixed reality data. To further improve our model and evaluate its performance on real data, we also provide a fully annotated real-world dataset with large environment and instance variation. Extensive experiments demonstrate that the proposed method is able to robustly estimate the pose and size of unseen object instances in real environments while also achieving state-of-the-art performance on standard 6D pose estimation benchmarks.