NOPE-SAC: Neural One-Plane RANSAC for Sparse-View Planar 3D Reconstruction

NOPE-SAC: Neural One-Plane RANSAC for Sparse-View Planar 3D Reconstruction
复制标题

DOI:
10.1109/tpami.2023.3314745
复制
发表时间:
2022-11
影响因子:
23.6
通讯作者:
Bin Tan;Nan Xue;Tianfu Wu;Guisong Xia
Bin Tan;Nan Xue;Tianfu Wu;Guisong Xia
中科院分区:
计算机科学1区
文献类型:
--
作者:
Bin Tan;Nan Xue;Tianfu Wu;Guisong Xia

文献摘要

相似文献

本文研究了在严格稀疏视图配置下具有挑战性的两视图三维重建问题,该问题受到输入图像对的对应性不足的困扰,无法用于相机姿态估计。我们提出了一种新的神经单平面RANSAC框架(简称NOPE-SAC),该框架利用神经网络从三维平面对应中学习单平面位姿假设的优异能力。在Siamese平面检测网络的基础上,我们的NOPE-SAC首先用粗糙的初始姿态生成假定的平面对应。然后将学习到的3D平面对应信息输入到共享mlp中,以估计单平面相机姿态假设,随后以RANSAC方式重新加权以获得最终相机姿态。由于神经单平面姿态最小化了自适应姿态假设生成的平面对应数量,因此对于稀疏视图输入,它可以使用少量平面对应实现稳定的姿态投票和可靠的姿态优化。在实验中,我们证明了我们的NOPE-SAC显著改善了具有严重视点变化的双视图输入的相机姿态估计,在两个具有挑战性的基准(即MatterPort3D和ScanNet)上设置了几个新的最先进的性能,用于稀疏视图3D重建。
This article studies the challenging two-view 3D reconstruction problem in a rigorous sparse-view configuration, which is suffering from insufficient correspondences in the input image pairs for camera pose estimation. We present a novel Neural One-PlanE RANSAC framework (termed NOPE-SAC in short) that exerts excellent capability of neural networks to learn one-plane pose hypotheses from 3D plane correspondences. Building on the top of a Siamese network for plane detection, our NOPE-SAC first generates putative plane correspondences with a coarse initial pose. It then feeds the learned 3D plane correspondences into shared MLPs to estimate the one-plane camera pose hypotheses, which are subsequently reweighed in a RANSAC manner to obtain the final camera pose. Because the neural one-plane pose minimizes the number of plane correspondences for adaptive pose hypotheses generation, it enables stable pose voting and reliable pose refinement with a few of plane correspondences for the sparse-view inputs. In the experiments, we demonstrate that our NOPE-SAC significantly improves the camera pose estimation for the two-view inputs with severe viewpoint changes, setting several new state-of-the-art performances on two challenging benchmarks, i.e., MatterPort3D and ScanNet, for sparse-view 3D reconstruction.