Multi-Level 3D CNN for Learning Multi-Scale Spatial Features

Multi-Level 3D CNN for Learning Multi-Scale Spatial Features
复制标题

DOI:
10.1109/cvprw.2019.00150
复制
发表时间:
2018-05
期刊:
2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
Sambit Ghadai;Xian Yeow Lee;Aditya Balu;S. Sarkar;A. Krishnamurthy
Sambit Ghadai;Xian Yeow Lee;Aditya Balu;S. Sarkar;A. Krishnamurthy
中科院分区:
其他
文献类型:
--
作者:
Sambit Ghadai;Xian Yeow Lee;Aditya Balu;S. Sarkar;A. Krishnamurthy

文献摘要

被引文献

相似文献

通过从对象的3D空间几何表示(诸如点云、3D模型、表面和RGB-D数据)学习多尺度空间特征,可以提高3D对象识别精度。目前的深度学习方法要么使用结构化数据表示(体素网格和八叉树),要么从非结构化表示(图形和点云)学习这些特征。从这种结构化表示中学习特征受到分辨率和树深度的限制,而非结构化表示由于数据样本之间的不均匀性而产生挑战。在本文中,我们提出了一个端到端的多级学习方法的多级体素网格,以克服这些缺点。为了证明所提出的多级学习的实用性,我们使用3D对象的多级体素表示来执行对象识别。多级体素表示由包含3D对象的体积信息的粗体素网格组成。此外,包含对象边界的一部分的粗网格中的每个体素被细分为多个精细级体素网格。我们的多级学习算法用于对象识别的性能与密集体素表示相当,同时使用显着更低的内存。
3D object recognition accuracy can be improved by learning the multi-scale spatial features from 3D spatial geometric representations of objects such as point clouds, 3D models, surfaces, and RGB-D data. Current deep learning approaches learn such features either using structured data representations (voxel grids and octrees) or from unstructured representations (graphs and point clouds). Learning features from such structured representations is limited by the restriction on resolution and tree depth while unstructured representations creates a challenge due to non-uniformity among data samples. In this paper, we propose an end-to-end multi-level learning approach on a multi-level voxel grid to overcome these drawbacks. To demonstrate the utility of the proposed multi-level learning, we use a multi-level voxel representation of 3D objects to perform object recognition. The multi-level voxel representation consists of a coarse voxel grid that contains volumetric information of the 3D object. In addition, each voxel in the coarse grid that contains a portion of the object boundary is subdivided into multiple fine-level voxel grids. The performance of our multi-level learning algorithm for object recognition is comparable to dense voxel representations while using significantly lower memory.