RGB-D Semantic Segmentation and Label-Oriented Voxelgrid Fusion for Accurate 3D Semantic Mapping

RGB-D Semantic Segmentation and Label-Oriented Voxelgrid Fusion for Accurate 3D Semantic Mapping
复制标题

RGB-D 语义分割和面向标签的体素网格融合,实现精确的 3D 语义映射

DOI:
10.1109/tcsvt.2021.3056726
复制
发表时间:
2022-01-01
影响因子:
8.4
通讯作者:
Zhang, Xiaolin
Zhang, Xiaolin
中科院分区:
工程技术1区
文献类型:
--
作者:
Shi, Wenjun;Xu, Jingwei;Zhang, Xiaolin

文献摘要

被引文献

相似文献

三维语义地图在各种应用中发挥着越来越重要的作用,特别是对于许多任务驱动的机器人。在本文中,我们提出了一种语义映射方法,从RGB-D扫描获得的三维语义地图。对比现有的方法,使用3D标注信息作为监督,我们专注于准确的2D帧标记和联合收割机标签在3D空间使用语义融合机制。对于场景解析,提出了一种具有新的歧视性掩模丢失的双流网络,以探索RGB和深度信息的充分提取和融合,实现稳定的语义分割。该判别掩模引导交叉熵损失函数并解释不同像素对反向传播的影响,从而减少深度噪声或对象边缘处的易出错注释的有害影响。在提供帧之间的对应关系后,这些语义帧被融合在统一的3D坐标使用新的面向标签的体素网格过滤器。该方法通过将面向标记的统计原则引入到标记点云中,保证了帧内的空间连续性和帧间的时空一致性。为了避免不相关帧之间的不利干扰,我们进一步提出了一种自适应分组算法,通过应用视锥过滤器分组帧有足够的重叠作为一个片段。为此,我们证明了所提出的方法的有效性的2D/3D语义标签基准的ScanNetv 2和Cityscapes数据集。
The 3D semantic map plays an increasingly important role in a wide variety of applications, especially for many kinds of task-driven robots. In this paper, we present a semantic mapping methodology for 3D semantic map obtaining from RGB-D scans. In contrast to existing methods that use 3D annotated information as supervisory, we focus on accurate 2D frame labeling and combine labels in 3D space using semantic fusion mechanism. For scene parsing, a two-stream network with a novel discriminatory mask loss is proposed to explore sufficient extraction and fusion of RGB and depth information achieving steadily semantic segmentation. The discriminatory mask guides the cross-entropy loss function and interprets the influence of different pixels on back-propagation, which reduces the harmful effects of the depth noise or the fallible annotation at the edges of objects. After the correspondences between frames are provided, these semantic frames are fused in unified 3D coordinates using the novel label-oriented voxelgrid filter. It can ensure the intra-frame spatial continuity and the inter-frame spatiotemporal consistency through introducing the label-oriented statistical principle into labeled point clouds. In order to avoid the unfavorable interference between uncorrelated frames, we further propose an adaptive grouping algorithm by applying the view frustum filter to group frames with sufficient overlap as a segment. To this end, we demonstrate the effectiveness of the proposed method on the 2D/3D semantic label benchmark of ScanNetv2 and Cityscapes datasets.