CVIDS: A Collaborative Localization and Dense Mapping Framework for Multi-Agent Based Visual-Inertial SLAM

CVIDS: A Collaborative Localization and Dense Mapping Framework for Multi-Agent Based Visual-Inertial SLAM
复制标题

DOI:
10.1109/tip.2022.3213189
复制
发表时间:
2022-10
影响因子:
10.6
通讯作者:
Tianjun Zhang;Lin Zhang;Yang Chen;Yicong Zhou
Tianjun Zhang;Lin Zhang;Yang Chen;Yicong Zhou
中科院分区:
计算机科学1区
文献类型:
--
作者:
Tianjun Zhang;Lin Zhang;Yang Chen;Yicong Zhou

文献摘要

相似文献

如今,视觉SLAM(同步定位与建图)因其成本低廉、应用范围广泛而成为研究热点。传统的视觉SLAM框架通常是为单代理系统设计的,通过单个机器人或移动设备上配备的传感器完成定位和建图。然而,单个代理人的机动性和工作能力通常是有限的。现实中,机器人或移动设备有时可能会以集群的形式部署,例如无人机编队、可穿戴动作捕捉系统等。据我们所知,现有的针对多智能体设计的SLAM系统仍然是零星的,并且大多数在功能上都存在着不可忽视的局限性。具体来说,一方面,大多数现有的多智能体SLAM系统只能提取一些关键特征并构建稀疏地图。另一方面,能够密集重建环境的方案无法摆脱对深度传感器的依赖,例如RGBD相机或LiDAR。暂时缺乏仅使用单目相机套件即可生成高密度地图的系统。为了在一定程度上填补研究空白​​,我们设计了一种新颖的协作SLAM系统,即CVIDS(协作视觉惯性密集SLAM),它遵循集中式松耦合框架,可以与任何现有的视觉惯性里程计(VIO)集成以完成共定位和密集重建。集成我们提出的鲁棒闭环检测模块和两阶段姿态图优化管道,CVIDS的共定位模块可以根据不同智能体客户端发送的打包图像和局部姿态,有效地估计统一坐标系中不同智能体的姿态。此外,我们基于运动的密集映射模块可以有效地恢复所选关键帧的 3D 结构,然后将其深度信息融合到全局映射中进行重建。 CVIDS 的优越性能得到了定量和定性实验结果的证实。为了使我们的结果具有可重复性,源代码已在 https://cslinzhang.github.io/CVIDS 上发布。
Nowadays, visual SLAM (Simultaneous Localization And Mapping) has become a hot research topic due to its low costs and wide application scopes. Traditional visual SLAM frameworks are usually designed for single-agent systems, completing both the localization and the mapping with sensors equipped on a single robot or a mobile device. However, the mobility and work capacity of the single agent are usually limited. In reality, robots or mobile devices sometimes may be deployed in the form of clusters, such as drone formations, wearable motion capture systems, and so on. As far as we know, existing SLAM systems designed for multi-agents are still sporadic, and most of them have non-negligible limitations in functions. Specifically, on one hand, most of the existing multi-agent SLAM systems can only extract some key features and build sparse maps. On the other hand, schemes that can reconstruct the environment densely cannot get rid of the dependence on depth sensors, such as RGBD cameras or LiDARs. Systems that can yield high-density maps just with monocular camera suites are temporarily lacking. As an attempt to fill in the research gap to some extent, we design a novel collaborative SLAM system, namely CVIDS (Collaborative Visual-Inertial Dense SLAM), which follows a centralized and loosely coupled framework and can be integrated with any existing Visual-Inertial Odometry (VIO) to accomplish the co-localization and the dense reconstruction. Integrating our proposed robust loop closure detection module and two-stage pose-graph optimization pipeline, the co-localization module of CVIDS can estimate the poses of different agents in a unified coordinate system efficiently from the packed images and local poses sent by the client-ends of different agents. Besides, our motion-based dense mapping module can effectively recover the 3D structures of selected keyframes and then fuse their depth information to the global map for reconstruction. The superior performance of CVIDS is corroborated by both quantitative and qualitative experimental results. To make our results reproducible, the source code has been released at https://cslinzhang.github.io/CVIDS.