NeRF-Pose: A First-Reconstruct-Then-Regress Approach for Weakly-supervised 6D Object Pose Estimation

NeRF-Pose: A First-Reconstruct-Then-Regress Approach for Weakly-supervised 6D Object Pose Estimation
复制标题

DOI:
10.1109/iccvw60793.2023.00226
复制
发表时间:
2022-03
期刊:
2023 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW)
影响因子:
--
通讯作者:
Fu Li;Hao Yu;I. Shugurov;Benjamin Busam;Shaowu Yang;Slobodan Ilic
Fu Li;Hao Yu;I. Shugurov;Benjamin Busam;Shaowu Yang;Slobodan Ilic
中科院分区:
其他
文献类型:
--
作者:
Fu Li;Hao Yu;I. Shugurov;Benjamin Busam;Shaowu Yang;Slobodan Ilic

文献摘要

被引文献

相似文献

单目图像中三维物体的位姿估计是计算机视觉中的一个基本而长期的问题。用于6D姿态估计的现有深度学习方法通常依赖于3D对象模型和6D姿态注释的可用性。然而,在真实的数据中的6D姿态的精确注释是复杂的、耗时的并且不可扩展的,而合成数据可扩展性好但缺乏真实性。为了避免这些问题,我们提出了一种基于弱监督重建的管道,称为NeRF-Pose,它在训练过程中只需要2D边界框和相对相机姿势。遵循先重建后回归的思想,我们首先以隐式神经表征的形式从多个视图重建对象。然后,我们训练一个姿态回归网络来预测图像和重建模型之间的像素级2D-3D对应关系。使用NeRF使能的PSENS +RANSAC算法来从预测的对应性估计稳定且准确的姿态。在LineMod和LineMod-Occlusion上的实验表明,与最好的6D姿态估计方法相比,尽管只使用弱标签进行训练,但所提出的方法具有最先进的精度。我们用真实的训练图像扩展了Homebrew DB数据集,以支持弱监督任务并取得令人信服的结果。扩展的数据集和代码将很快发布。
Pose estimation of 3D objects in monocular images is a fundamental and long-standing problem in computer vision. Existing deep learning approaches for 6D pose estimation typically rely on the availability of 3D object models and 6D pose annotations. However, precise annotation of 6D poses in real data is intricate, time-consuming and not scalable, while synthetic data scales well but lacks realism. To avoid these problems, we present a weakly-supervised reconstruction-based pipeline, named NeRF-Pose, which needs only 2D bounding boxes and relative camera poses during training. Following the first-reconstruct-then-regress idea, we first reconstruct the objects from multiple views in the form of an implicit neural representation. Then, we train a pose regression network to predict pixel-wise 2D-3D correspondences between images and the reconstructed model. A NeRF-enabled PnP+RANSAC algorithm is used to estimate stable and accurate pose from the predicted correspondences. Experiments on LineMod and LineMod-Occlusion show that the proposed method has state-of-the-art accuracy in comparison to the best 6D pose estimation methods in spite of being trained only with weak labels. We extend the Homebrewed DB dataset with real training images to support the weakly supervised task and achieve compelling results. The extended dataset and code will be released soon.