Robust Category-Level 6D Pose Estimation with Coarse-to-Fine Rendering of Neural Features

Robust Category-Level 6D Pose Estimation with Coarse-to-Fine Rendering of Neural Features
复制标题

DOI:
10.48550/arxiv.2209.05624
复制
发表时间:
2022-09
期刊:
--
影响因子:
--
通讯作者:
Wufei Ma;Angtian Wang;A. Yuille;Adam Kortylewski
Wufei Ma;Angtian Wang;A. Yuille;Adam Kortylewski
中科院分区:
其他
文献类型:
--
作者:
Wufei Ma;Angtian Wang;A. Yuille;Adam Kortylewski

文献摘要

相似文献

我们考虑从单个 RGB 图像进行类别级 6D 姿态估计的问题。我们的方法将对象类别表示为长方体网格,并学习每个网格顶点的神经特征激活的生成模型,以通过可微渲染执行姿势估计。基于渲染的方法的一个常见问题是它们依赖于边界框建议,而边界框建议不传达有关对象 3D 旋转的信息,并且当对象部分被遮挡时并不可靠。相反,我们引入了一种从粗到细的优化策略,该策略利用渲染过程来估计一组稀疏的 6D 对象建议,随后通过基于梯度的优化进行细化。实现我们方法收敛的关键是使用对比学习将神经特征表示训练为尺度和旋转不变。我们的实验证明了与之前的工作相比,类别级 6D 姿态估计性能得到了增强,特别是在强部分遮挡的情况下。
We consider the problem of category-level 6D pose estimation from a single RGB image. Our approach represents an object category as a cuboid mesh and learns a generative model of the neural feature activations at each mesh vertex to perform pose estimation through differentiable rendering. A common problem of rendering-based approaches is that they rely on bounding box proposals, which do not convey information about the 3D rotation of the object and are not reliable when objects are partially occluded. Instead, we introduce a coarse-to-fine optimization strategy that utilizes the rendering process to estimate a sparse set of 6D object proposals, which are subsequently refined with gradient-based optimization. The key to enabling the convergence of our approach is a neural feature representation that is trained to be scale- and rotation-invariant using contrastive learning. Our experiments demonstrate an enhanced category-level 6D pose estimation performance compared to prior work, particularly under strong partial occlusion.