Lossy Point Cloud Geometry Compression via End-to-End Learning

Lossy Point Cloud Geometry Compression via End-to-End Learning
复制标题

DOI:
10.1109/tcsvt.2021.3051377
复制
发表时间:
2019-09
影响因子:
8.4
通讯作者:
Jianqiang Wang;Hao Zhu;Haojie Liu-;Zhan Ma
Jianqiang Wang;Hao Zhu;Haojie Liu-;Zhan Ma
中科院分区:
工程技术1区
文献类型:
--
作者:
Jianqiang Wang;Hao Zhu;Haojie Liu-;Zhan Ma

文献摘要

被引文献

相似文献

提出了一种新的端到端学习型点云几何压缩(又称学习型点云几何压缩)系统,该系统利用基于堆叠深度神经网络(DNN)的变分自动编码器(VAE)来有效地压缩点云几何(PCG)。在这种系统的探索中,首先对PCG进行体素化,并将其划分为不重叠的3D立方体,然后将这些立方体送入堆叠的3D卷积中,以获得紧凑的潜在特征和超先验生成。利用超先验知识改进了对经熵编码的潜在特征的条件概率建模。在训练中采用加权的二进制交叉熵(WBCE)损失,在推理中使用自适应阈值来去除虚假体素,减少失真。客观上,我们的方法超过了运动图像专家组(MPEG)标准化的基于几何的点云压缩(G-PCC)算法,具有显著的性能差距,例如,使用公共测试数据集和其他公共数据集,至少节省了60%的BD率(Bjöntegaard Delta Rate)。主观上,与所有现有的符合MPEG标准的PCC方法相比,我们的方法具有更好的视觉质量、更平滑的表面重建和更吸引人的细节。我们的方法总共需要大约2.5MB的参数,这对于实际实现来说是一个相当小的尺寸,即使在嵌入式平台上也是如此。其他消融研究分析了各种方面(如阈值、核等),以检查我们的学习-PCGC的泛化和应用能力。我们希望在https://njuvision.github.io/PCGCv1/上公开所有材料,以便进行可重现的研究。
This paper presents a novel end-to-end Learned Point Cloud Geometry Compression (a.k.a., Learned-PCGC) system, leveraging stacked Deep Neural Networks (DNN) based Variational AutoEncoder (VAE) to efficiently compress the Point Cloud Geometry (PCG). In this systematic exploration, PCG is first voxelized, and partitioned into non-overlapped 3D cubes, which are then fed into stacked 3D convolutions for compact latent feature and hyperprior generation. Hyperpriors are used to improve the conditional probability modeling of entropy-coded latent features. A Weighted Binary Cross-Entropy (WBCE) loss is applied in training while an adaptive thresholding is used in inference to remove false voxels and reduce the distortion. Objectively, our method exceeds the Geometry-based Point Cloud Compression (G-PCC) algorithm standardized by the Moving Picture Experts Group (MPEG) with a significant performance margin, e.g., at least 60% BD-Rate (Bjöntegaard Delta Rate) savings, using common test datasets, and other public datasets. Subjectively, our method has presented better visual quality with smoother surface reconstruction and appealing details, in comparison to all existing MPEG standard compliant PCC methods. Our method requires about 2.5 MB parameters in total, which is a fairly small size for practical implementation, even on embedded platform. Additional ablation studies analyze a variety of aspects (e.g., thresholding, kernels, etc) to examine the generalization, and application capacity of our Learned-PCGC. We would like to make all materials publicly accessible at https://njuvision.github.io/PCGCv1/ for reproducible research.