Unsupervised Scale-Consistent Depth Learning from Video

Unsupervised Scale-Consistent Depth Learning from Video
复制标题

DOI:
10.1007/s11263-021-01484-6
复制
发表时间:
2021-06-18
影响因子:
19.5
通讯作者:
Reid, Ian
Reid, Ian
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bian, Jia-Wang;Zhan, Huangying;Reid, Ian

文献摘要

被引文献

相似文献

我们提出了一种单目深度估计方法SC-Depth,该方法只需要训练未标记的视频,并且能够在推理时进行尺度一致的预测。我们的贡献包括:(I)我们提出了几何一致性损失,它惩罚了相邻视图之间预测深度的不一致;(Ii)我们提出了一种自发现掩模来自动定位违反基本静态场景假设并在训练过程中产生噪声信号的运动目标;(Iii)我们通过详细的烧蚀研究证明了每个分量的有效性,并在Kitti和NYUv2数据集上显示了高质量的深度估计结果。此外,由于具有尺度一致性预测能力,我们表明我们的单目训练的深度网络可以很容易地集成到ORB-SLAM2系统中,以实现更稳健和准确的跟踪。所提出的混合伪RGBD SLAM在Kitti上显示了令人信服的结果,并且它在不需要额外训练的情况下很好地推广到KAIST数据集。最后,我们提供了几个用于定性评估的示例。源代码在GitHub上发布。
We propose a monocular depth estimation method SC-Depth, which requires only unlabelled videos for training and enables the scale-consistent prediction at inference time. Our contributions include: (i) we propose a geometry consistency loss, which penalizes the inconsistency of predicted depths between adjacent views; (ii) we propose a self-discovered mask to automatically localize moving objects that violate the underlying static scene assumption and cause noisy signals during training; (iii) we demonstrate the efficacy of each component with a detailed ablation study and show high-quality depth estimation results in both KITTI and NYUv2 datasets. Moreover, thanks to the capability of scale-consistent prediction, we show that our monocular-trained deep networks are readily integrated into ORB-SLAM2 system for more robust and accurate tracking. The proposed hybrid Pseudo-RGBD SLAM shows compelling results in KITTI, and it generalizes well to the KAIST dataset without additional training. Finally, we provide several demos for qualitative evaluation. The source code is released on GitHub.