Deep Non-Rigid Structure From Motion With Missing Data

Deep Non-Rigid Structure From Motion With Missing Data
复制标题

DOI:
10.1109/tpami.2020.2997026
复制
发表时间:
2019-07
影响因子:
23.6
通讯作者:
Chen Kong;S. Lucey
Chen Kong;S. Lucey
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chen Kong;S. Lucey

文献摘要

相似文献

非刚体运动结构(NRSfM)是指从具有二维对应关系的图像集合中重建相机和非刚体物体的三维点云的问题。目前的NRSfM算法在两个方面受到限制:(i)图像的数量,(ii)它们可以处理的形状变化的类型。这些困难源于系统条件和需要建模的自由度之间的内在冲突——这阻碍了它在许多视觉应用中的实际效用。在本文中,我们为NRSFM提出了一种新的分层稀疏编码模型,该模型可以克服(i)和(ii),以至于NRSFM可以应用于以前认为过于病态的视觉问题。我们的方法在实践中是通过训练一个无监督深度神经网络(DNN)自编码器来实现的,该编码器具有独特的结构,能够从3D结构中分离出姿态。使用现代深度学习计算平台使我们能够以前所未有的规模和形状复杂性解决NRSfM问题。我们的方法没有3D监督,仅依赖于2D点对应。此外,我们的方法还能够处理缺失/遮挡的2D点,而不需要矩阵补全。大量的实验证明了我们的方法令人印象深刻的性能,在某些情况下,我们对所有现有的最先进的作品表现出了卓越的精度和鲁棒性。我们进一步提出了一种新的质量度量(基于网络权重),它绕过了对三维地面真实度的需要,以确定我们对可重构性的信心。我们相信,我们的工作是一个重大的进步,在最先进的NRSFM。
Non-rigid structure from motion (NRSfM) refers to the problem of reconstructing cameras and the 3D point cloud of a non-rigid object from an ensemble of images with 2D correspondences. Current NRSfM algorithms are limited from two perspectives: (i) the number of images, and (ii) the type of shape variability they can handle. These difficulties stem from the inherent conflict between the condition of the system and the degrees of freedom needing to be modeled – which has hampered its practical utility for many applications within vision. In this paper we propose a novel hierarchical sparse coding model for NRSFM which can overcome (i) and (ii) to such an extent, that NRSFM can be applied to problems in vision previously thought too ill posed. Our approach is realized in practice as the training of an unsupervised deep neural network (DNN) auto-encoder with a unique architecture that is able to disentangle pose from 3D structure. Using modern deep learning computational platforms allows us to solve NRSfM problems at an unprecedented scale and shape complexity. Our approach has no 3D supervision, relying solely on 2D point correspondences. Further, our approach is also able to handle missing/occluded 2D points without the need for matrix completion. Extensive experiments demonstrate the impressive performance of our approach where we exhibit superior precision and robustness against all available state-of-the-art works in some instances by an order of magnitude. We further propose a new quality measure (based on the network weights) which circumvents the need for 3D ground-truth to ascertain the confidence we have in the reconstructability. We believe our work to be a significant advance over state-of-the-art in NRSFM.