Towards the Probabilistic Fusion of Learned Priors into Standard Pipelines for 3D Reconstruction

Towards the Probabilistic Fusion of Learned Priors into Standard Pipelines for 3D Reconstruction
复制标题

DOI:
10.1109/icra40945.2020.9197001
复制
发表时间:
2020-05
期刊:
2020 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Tristan Laidlow;J. Czarnowski;Andrea Nicastro;R. Clark;Stefan Leutenegger
Tristan Laidlow;J. Czarnowski;Andrea Nicastro;R. Clark;Stefan Leutenegger
中科院分区:
其他
文献类型:
--
作者:
Tristan Laidlow;J. Czarnowski;Andrea Nicastro;R. Clark;Stefan Leutenegger

文献摘要

相似文献

将深度学习的结果与标准3D重建管道联合收割机结合的最佳方法仍然是一个悬而未决的问题。虽然将传统多视图立体方法的输出传递到网络以进行正则化或细化的系统目前似乎获得最佳结果,但可能优选将深度神经网络视为单独的组件,其结果可以概率地融合到基于几何的系统中。不幸的是,进行这种类型的融合所需的误差模型还没有得到很好的理解,提出了许多不同的方法。最近,一些系统通过使其网络预测概率分布而不是单个值来实现良好的结果。我们建议使用这种方法将学习到的单视图深度先验融合到标准的3D重建系统中。我们的系统能够为一组关键帧增量地生成密集的深度图。我们训练一个深度神经网络来预测单个图像中每个像素深度的离散非参数概率分布。然后,我们将此“概率体积”与基于后续帧和关键帧图像之间的光度一致性的另一个概率体积融合。我们认为,结合这两个来源的概率卷将导致更好的条件卷。为了从体积中提取深度图,我们最小化成本函数,该成本函数包括基于网络预测的表面法线和遮挡边界的正则化项。通过一系列的实验,我们证明,这些组件中的每一个提高了系统的整体性能。
The best way to combine the results of deep learning with standard 3D reconstruction pipelines remains an open problem. While systems that pass the output of traditional multi-view stereo approaches to a network for regularisation or refinement currently seem to get the best results, it may be preferable to treat deep neural networks as separate components whose results can be probabilistically fused into geometry- based systems. Unfortunately, the error models required to do this type of fusion are not well understood, with many different approaches being put forward. Recently, a few systems have achieved good results by having their networks predict probability distributions rather than single values. We propose using this approach to fuse a learned single-view depth prior into a standard 3D reconstruction system.Our system is capable of incrementally producing dense depth maps for a set of keyframes. We train a deep neural network to predict discrete, nonparametric probability distributions for the depth of each pixel from a single image. We then fuse this "probability volume" with another probability volume based on the photometric consistency between subsequent frames and the keyframe image. We argue that combining the probability volumes from these two sources will result in a volume that is better conditioned. To extract depth maps from the volume, we minimise a cost function that includes a regularisation term based on network predicted surface normals and occlusion boundaries. Through a series of experiments, we demonstrate that each of these components improves the overall performance of the system.