Real-Time Dense Monocular SLAM With Online Adapted Depth Prediction Network

Real-Time Dense Monocular SLAM With Online Adapted Depth Prediction Network
复制标题

DOI:
10.1109/tmm.2018.2859034
复制
发表时间:
2019-02
影响因子:
7.3
通讯作者:
Hongcheng Luo;Yang Gao;Yuhao Wu;Chunyuan Liao;Xin Yang;Kwang-Ting Cheng
Hongcheng Luo;Yang Gao;Yuhao Wu;Chunyuan Liao;Xin Yang;Kwang-Ting Cheng
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hongcheng Luo;Yang Gao;Yuhao Wu;Chunyuan Liao;Xin Yang;Kwang-Ting Cheng

文献摘要

被引文献

相似文献

在过去的几年里,通过卷积神经网络(CNN)从单个图像估计深度图方面已经取得了相当大的进步。将 CNN 的深度预测与传统的单目同步定位与建图 (SLAM) 相结合,有望实现准确而密集的单目重建,特别是解决传统单目 SLAM 中长期存在的两个挑战:地图完整性低和尺度模糊。然而,预训练的 CNN 估计的深度通常无法从训练数据中获得足够的精度来应对不同类型的环境,这对于某些应用(例如无人机在未知场景中避障)很常见。此外,CNN 的深度预测不准确可能会在单目 SLAM 中产生较大的跟踪误差。在本文中,我们提出了一种实时密集单目 SLAM 系统,该系统有效地将直接单目 SLAM 与在线自适应深度预测网络融合,以实现根据训练数据对不同类型场景进行准确的深度预测,并为跟踪和建图提供绝对尺度信息。具体来说,一方面,直接 SLAM 的跟踪姿势(即平移和旋转)用于选择一小组高效且可靠的训练图像,这些图像充当实时调整深度预测网络的地面实况,以提高对不同类型场景的泛化能力。引入了具有选择性更新策略的分阶段随机梯度下降算法,以实现调整过程的有效收敛。另一方面,自适应网络生成的密集地图被应用于解决直接单目 SLAM 的尺度模糊性,从而提高了跟踪和整体重建的准确性。该系统在CPU和GPU的辅助下,可以实现实时性能,并逐步提高重建精度。公共数据集的实验结果和无人机避障的实时应用表明,我们的方法优于最先进的方法,具有更高的地图完整性和准确性,以及更小的跟踪误差。
Considerable advances have been achieved in estimating the depth map from a single image via convolutional neural networks (CNNs) during the past few years. Combining depth prediction from CNNs with conventional monocular simultaneous localization and mapping (SLAM) is promising for accurate and dense monocular reconstruction, in particular addressing the two long-standing challenges in conventional monocular SLAM: low map completeness and scale ambiguity. However, depth estimated by pretrained CNNs usually fails to achieve sufficient accuracy for environments of different types from the training data, which are common for certain applications such as obstacle avoidance of drones in unknown scenes. Additionally, inaccurate depth prediction of CNN could yield large tracking errors in monocular SLAM. In this paper, we present a real-time dense monocular SLAM system, which effectively fuses direct monocular SLAM with an online-adapted depth prediction network for achieving accurate depth prediction of scenes of different types from the training data and providing absolute scale information for tracking and mapping. Specifically, on one hand, tracking pose (i.e., translation and rotation) from direct SLAM is used for selecting a small set of highly effective and reliable training images, which acts as ground truth for tuning the depth prediction network on-the-fly toward better generalization ability for scenes of different types. A stage-wise Stochastic Gradient Descent algorithm with a selective update strategy is introduced for efficient convergence of the tuning process. On the other hand, the dense map produced by the adapted network is applied to address scale ambiguity of direct monocular SLAM which in turn improves the accuracy of both tracking and overall reconstruction. The system with assistance of both CPUs and GPUs, can achieve real-time performance with progressively improved reconstruction accuracy. Experimental results on public datasets and live application to obstacle avoidance of drones demonstrate that our method outperforms the state-of-the-art methods with greater map completeness and accuracy, and a smaller tracking error.