MOTSLAM: MOT-assisted monocular dynamic SLAM using single-view depth estimation

MOTSLAM: MOT-assisted monocular dynamic SLAM using single-view depth estimation
复制标题

DOI:
10.1109/iros47612.2022.9982280
复制
发表时间:
2022-10
期刊:
2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Hanwei Zhang;Hideaki Uchiyama;Shintaro Ono;Hiroshi Kawasaki
Hanwei Zhang;Hideaki Uchiyama;Shintaro Ono;Hiroshi Kawasaki
中科院分区:
其他
文献类型:
--
作者:
Hanwei Zhang;Hideaki Uchiyama;Shintaro Ono;Hiroshi Kawasaki

文献摘要

相似文献

针对静态场景的视觉SLAM系统已经开发出来,具有令人满意的精度和鲁棒性。随着自动驾驶、增强现实和虚拟现实等各种场景对动态环境的理解,动态3D目标跟踪已经成为视觉SLAM中的一项重要能力。然而,仅利用单目图像进行动态SLAM仍然是一个具有挑战性的问题,因为很难关联动态特征并估计它们的位置。本文提出了一种单目结构的动态视觉SLAM系统MOTSLAM,它可以同时跟踪动态物体的姿态和包围盒。MOTSLAM首先执行具有关联的2D和3D边界框检测的多对象跟踪(MOT),以创建初始3D对象。然后,采用基于神经网络的单目深度估计方法提取动态特征的深度。最后,相机姿势、对象姿势以及静态和动态地图点都使用了一种新颖的捆绑调整进行了联合优化。我们在Kitti数据集上的实验表明,我们的系统在摄像机自我运动和单目动态SLAM上的目标跟踪方面都达到了最好的性能。
Visual SLAM systems targeting static scenes have been developed with satisfactory accuracy and robustness. Dynamic 3D object tracking has then become a significant capability in visual SLAM with the requirement of under-standing dynamic surroundings in various scenarios including autonomous driving, augmented and virtual reality. However, performing dynamic SLAM solely with monocular images remains a challenging problem due to the difficulty of asso-ciating dynamic features and estimating their positions. In this paper, we present MOTSLAM, a dynamic visual SLAM system with the monocular configuration that tracks both poses and bounding boxes of dynamic objects. MOTSLAM first performs multiple object tracking (MOT) with associated both 2D and 3D bounding box detection to create initial 3D objects. Then, neural-network-based monocular depth estimation is applied to fetch the depth of dynamic features. Finally, camera poses, object poses, and both static, as well as dynamic map points, are jointly optimized using a novel bundle adjustment. Our experiments on the KITTI dataset demonstrate that our system has reached best performance on both camera ego-motion and object tracking on monocular dynamic SLAM.