Voxel-Based Representation Learning for Place Recognition Based on 3D Point Clouds

Voxel-Based Representation Learning for Place Recognition Based on 3D Point Clouds
复制标题

DOI:
10.1109/iros45743.2020.9340992
复制
发表时间:
2020-10
期刊:
2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
S. Siva;Zachary Nahman;Hao Zhang
S. Siva;Zachary Nahman;Hao Zhang
中科院分区:
其他
文献类型:
--
作者:
S. Siva;Zachary Nahman;Hao Zhang

文献摘要

相似文献

地点识别是解决同步定位与建图 (SLAM) 关键问题的关键组成部分。大多数现有方法使用视觉图像;然而,使用 3D 点云的位置识别,尤其是基于体素表示的位置识别,尚未得到很好的解决。在本文中,我们介绍了基于体素的表示学习 (VBRL) 的新方法,该方法使用 3D 点云来识别具有长期环境变化的位置。 VBRL 将 3D 点云输入分割为体素,并使用从这些体素中提取的多模态特征来执行地点识别。此外,VBRL 使用结构化稀疏诱导规范来学习代表性体素和特征模态,这对于匹配长期变化下的位置非常重要。地点识别、体素和特征学习都集成到统一的正则化优化公式中。由于稀疏性范数是非光滑的,因此很难解决公式化的优化问题。因此,我们设计了一种新的迭代优化算法,该算法具有理论上的收敛性保证。实验结果表明,VBRL 使用 3D 点云数据可以很好地进行地点识别,并且能够学习体素和特征模态的重要性。
Place recognition is a critical component towards addressing the key problem of Simultaneous Localization and Mapping (SLAM). Most existing methods use visual images; whereas, place recognition using 3D point clouds, especially based on the voxel representations, has not been well addressed yet. In this paper, we introduce the novel approach of voxel-based representation learning (VBRL) that uses 3D point clouds to recognize places with long-term environment variations. VBRL splits a 3D point cloud input into voxels and uses multi-modal features extracted from these voxels to perform place recognition. Additionally, VBRL uses structured sparsity-inducing norms to learn representative voxels and feature modalities that are important to match places under long-term changes. Both place recognition, and voxel and feature learning are integrated into a unified regularized optimization formulation. As the sparsity-inducing norms are non-smooth, it is hard to solve the formulated optimization problem. Thus, we design a new iterative optimization algorithm, which has a theoretical convergence guarantee. Experimental results have shown that VBRL performs place recognition well using 3D point cloud data and is capable of learning the importance of voxels and feature modalities.