MDN-VO: Estimating Visual Odometry with Confidence

MDN-VO: Estimating Visual Odometry with Confidence
复制标题

DOI:
10.1109/iros51168.2021.9636827
复制
发表时间:
2021-09
期刊:
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Nimet Kaygusuz;Oscar Alejandro Mendez Maldonado;R. Bowden
Nimet Kaygusuz;Oscar Alejandro Mendez Maldonado;R. Bowden
中科院分区:
其他
文献类型:
--
作者:
Nimet Kaygusuz;Oscar Alejandro Mendez Maldonado;R. Bowden

文献摘要

相似文献

视觉里程计(VO)用于许多应用,包括机器人和自主系统。然而,传统的基于特征匹配的方法计算量大,不能直接解决故障情况,而是依赖于启发式方法来检测故障。在这项工作中,我们提出了一个基于深度学习的VO模型来有效地估计6-DoF姿势,以及这些估计的置信度模型。我们利用CNN - RNN混合模型从图像序列中学习特征表示。然后,我们采用混合密度网络(MDN)估计相机运动的高斯混合物,基于提取的时空表示。我们的模型使用姿势标签作为监督的来源,但以无监督的方式获得不确定性。我们在KITTI和nuScenes数据集上评估了所提出的模型,并报告了大量的定量和定性结果,以分析姿态和不确定性估计的性能。我们的实验表明,该模型超过了国家的最先进的性能,除了使用预测的姿势不确定性检测故障情况。
Visual Odometry (VO) is used in many applications including robotics and autonomous systems. However, traditional approaches based on feature matching are computationally expensive and do not directly address failure cases, instead relying on heuristic methods to detect failure. In this work, we propose a deep learning-based VO model to efficiently estimate 6-DoF poses, as well as a confidence model for these estimates. We utilise a CNN - RNN hybrid model to learn feature representations from image sequences. We then employ a Mixture Density Network (MDN) which estimates camera motion as a mixture of Gaussians, based on the extracted spatio-temporal representations. Our model uses pose labels as a source of supervision, but derives uncertainties in an unsupervised manner. We evaluate the proposed model on the KITTI and nuScenes datasets and report extensive quantitative and qualitative results to analyse the performance of both pose and uncertainty estimation. Our experiments show that the proposed model exceeds state-of-the-art performance in addition to detecting failure cases using the predicted pose uncertainty.