ViewNet: A Novel Projection-Based Backbone with View Pooling for Few-shot Point Cloud Classification

ViewNet: A Novel Projection-Based Backbone with View Pooling for Few-shot Point Cloud Classification
复制标题

DOI:
10.1109/cvpr52729.2023.01693
复制
发表时间:
2023-06
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Jiajing Chen;Min Yang;Senem Velipasalar
Jiajing Chen;Min Yang;Senem Velipasalar
中科院分区:
其他
文献类型:
--
作者:
Jiajing Chen;Min Yang;Senem Velipasalar

文献摘要

相似文献

尽管针对 3D 点云相关任务提出了不同的方法,但 3D 点云的少样本学习 (FSL) 仍处于探索之中。在 FSL 中,与传统的监督学习不同,训练数据和测试数据的类别不重叠,并且模型只需要从少数样本中识别不可见的类别。现有的 3D 点云 FSL 方法采用基于点的模型作为其骨干。然而,基于我们广泛的实验和分析,我们首先表明使用基于点的主干并不是最合适的 FSL 方法,因为 (i) 3D 基于点的主干中使用的最大池化操作丢弃了大量点的特征,降低了表示形状信息的能力; (ii) 基于点的主干对遮挡很敏感。为了解决这些问题,我们建议采用基于投影和 2D 卷积神经网络的主干网(称为 ViewNet),用于 3D 点云的 FSL。我们的方法首先将 3D 点云投影到六个不同的视图上,以缓解丢失点的问题。此外,为了生成更具描述性和区分性的特征,我们提出了视图池化,它将不同的投影平面组合组合成五组,并对每组执行最大池化。在 ModelNet40、ScanObjectNN 和 ModelNet40-C 数据集上进行的实验以及交叉验证表明,我们的方法始终优于最先进的基线。此外,与传统的图像分类主干网(例如 ResNet)相比,所提出的 ViewNet 可以从点云的多个视图中提取更多的区分特征。我们还表明,ViewNet 可以用作具有不同 FSL 头的主干网,并且与传统使用的主干网相比,提供了改进的性能。
Although different approaches have been proposed for 3D point cloud-related tasks, few-shot learning (FSL) of 3D point clouds still remains under-explored. In FSL, un-like traditional supervised learning, the classes of training and test data do not overlap, and a model needs to rec-ognize unseen classes from only a few samples. Existing FSL methods for 3D point clouds employ point-based models as their backbone. Yet, based on our extensive experiments and analysis, we first show that using a point-based backbone is not the most suitable FSL approach, since (i) a large number of points' features are discarded by the max pooling operation used in 3D point-based backbones, decreasing the ability of representing shape information; (ii) point-based backbones are sensitive to occlusion. To address these issues, we propose employing a projection-and 2D Convolutional Neural Network-based backbone, referred to as the ViewNet, for FSL from 3D point clouds. Our approach first projects a 3D point cloud onto six different views to alleviate the issue of missing points. Also, to generate more descriptive and distinguishing features, we propose View Pooling, which combines different projected plane combinations into five groups and performs max-pooling on each of them. The experiments performed on the ModelNet40, ScanObjectNN and ModelNet40-C datasets, with cross validation, show that our method consistently outperforms the state-of-the-art baselines. Moreover, compared to traditional image classification backbones, such as ResNet, the proposed ViewNet can extract more distinguishing features from multiple views of a point cloud. We also show that ViewNet can be used as a backbone with different FSL heads and provides improved performance compared to traditionally used backbones.