RGAM: A novel network architecture for 3D point cloud semantic segmentation in indoor scenes

RGAM: A novel network architecture for 3D point cloud semantic segmentation in indoor scenes
复制标题

DOI:
10.1016/j.ins.2021.04.069
复制
发表时间:
2021-05-06
影响因子:
8.1
通讯作者:
Wang, Rui
Wang, Rui
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chen, Xue-Tao;Li, Ying;Wang, Rui

文献摘要

被引文献

相似文献

三维点云语义分割是计算机视觉场景理解的重要组成部分。然而,由于缺乏细节,现有的网络缺乏识别复杂场景的能力。本文提出了一种新的网络结构,称为带注意模块的环分组神经网络(RGAM),它在现有网络的基础上进行了四个改进。首先,设计了新颖的多尺度环分组学习,在不重叠采样的情况下提取多尺度邻域特征,使网络能够适应不同尺度的目标;其次,将邻域信息融合定义为多个邻域特征的加权和,使每个点的表示能够在不同的邻域中得到考虑。第三,在全局视图中,在邻域之间引入空间关注模块,允许利用远程上下文信息进行三维点云语义分割。最后,在RGAM的基础上增加通道关注模块,各通道与关键信息的相关性增强了RGAM对复杂场景的识别能力。在具有挑战性的S3DIS、ScanNet和NYU-V2数据集上的实验结果表明,RGAM在3D点云语义分割方面比基于几种最先进算法的现有网络具有更强的识别能力。(c) 2021爱思唯尔公司版权所有。
Three-dimensional (3D) point cloud semantic segmentation is an essential part of computer vision for scene comprehension. Nevertheless, due to their loss of detail, existing networks lack the ability to recognize complex scenes. This paper proposes a novel network architecture, called the ring grouping neural network with attention module (RGAM), which presents four improvements over the existing networks. First, novel multi-scale ring grouping learning is designed to extract the multi-scale neighborhood features without overlapped sampling, allowing the network to adapt to objects of different scales. Second, neighborhood information fusion is defined as the weighted sum of multiple neighborhood features, enabling the representation of each point to be considered in different neighborhoods. Third, in the global view, a spatial attention module is introduced among the neighborhoods, allowing long-range contextual information to be exploited for 3D point cloud semantic segmentation. Finally, a channel attention module is appended to the RGAM: the correlation of each channel with key information enhances the complex scene recognition ability of the RGAM. Experimental results on the challenging S3DIS, ScanNet, and NYU-V2 datasets demonstrate that the RGAM has stronger recognition ability than the existing networks based on several state-of-the-art algorithms for 3D point cloud semantic segmentation. (c) 2021 Elsevier Inc. All rights reserved.