IV-SLAM: Introspective Vision for Simultaneous Localization and Mapping

IV-SLAM: Introspective Vision for Simultaneous Localization and Mapping
复制标题

DOI:
--
复制
发表时间:
2020-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Sadegh Rabiee;Joydeep Biswas
Sadegh Rabiee;Joydeep Biswas
中科院分区:
其他
文献类型:
--
作者:
Sadegh Rabiee;Joydeep Biswas

文献摘要

被引文献

相似文献

视觉同步定位和映射(V-SLAM)的现有解决方案假设特征提取和匹配中的误差是独立且同分布的(i.i.d),但是已知该假设不是真的-从图像的低对比度区域提取的特征比来自尖角的特征表现出更宽的误差分布。此外,当感测到的图像包括诸如镜面反射、透镜光斑或动态对象的阴影之类的挑战性条件时,V-SLAM算法易于发生灾难性的跟踪故障。为了解决这些问题,以前的工作集中在构建更强大的视觉前端,以过滤掉具有挑战性的功能。在本文中,我们提出了SLAM(IV-SLAM)的内省愿景,这是一种解决这些挑战的根本不同的方法。IV-SLAM明确地将来自视觉特征的重投影误差的噪声过程建模为上下文相关的,并且因此非独立同分布。我们引入了一种自主监督的方法,用于IV-SLAM收集训练数据来学习这种上下文感知的噪声模型。使用该学习的噪声模型,IV-SLAM引导特征提取以从图像的可能导致较低噪声的部分中选择更多特征,并且进一步将学习的噪声模型并入联合最大似然估计中,从而使其对上述类型的错误具有鲁棒性。我们提出的实证结果表明,IV-SLAM 1)能够准确地预测输入图像中的误差来源,2)与V-SLAM相比,减少了跟踪误差,3)与V-SLAM相比,在具有挑战性的真实的机器人数据上,跟踪失败之间的平均距离增加了70%以上。
Existing solutions to visual simultaneous localization and mapping (V-SLAM) assume that errors in feature extraction and matching are independent and identically distributed (i.i.d), but this assumption is known to not be true -- features extracted from low-contrast regions of images exhibit wider error distributions than features from sharp corners. Furthermore, V-SLAM algorithms are prone to catastrophic tracking failures when sensed images include challenging conditions such as specular reflections, lens flare, or shadows of dynamic objects. To address such failures, previous work has focused on building more robust visual frontends, to filter out challenging features. In this paper, we present introspective vision for SLAM (IV-SLAM), a fundamentally different approach for addressing these challenges. IV-SLAM explicitly models the noise process of reprojection errors from visual features to be context-dependent, and hence non-i.i.d. We introduce an autonomously supervised approach for IV-SLAM to collect training data to learn such a context-aware noise model. Using this learned noise model, IV-SLAM guides feature extraction to select more features from parts of the image that are likely to result in lower noise, and further incorporate the learned noise model into the joint maximum likelihood estimation, thus making it robust to the aforementioned types of errors. We present empirical results to demonstrate that IV-SLAM 1) is able to accurately predict sources of error in input images, 2) reduces tracking error compared to V-SLAM, and 3) increases the mean distance between tracking failures by more than 70% on challenging real robot data compared to V-SLAM.