Shape My Face: Registering 3D Face Scans by Surface-to-Surface Translation

Shape My Face: Registering 3D Face Scans by Surface-to-Surface Translation
复制标题

DOI:
10.1007/s11263-021-01494-4
复制
发表时间:
2020-12
影响因子:
19.5
通讯作者:
Mehdi Bahri;Eimear O' Sullivan;Shunwang Gong;Feng Liu;Xiaoming Liu;M. Bronstein;S. Zafeiriou
Mehdi Bahri;Eimear O' Sullivan;Shunwang Gong;Feng Liu;Xiaoming Liu;M. Bronstein;S. Zafeiriou
中科院分区:
计算机科学2区
文献类型:
--
作者:
Mehdi Bahri;Eimear O' Sullivan;Shunwang Gong;Feng Liu;Xiaoming Liu;M. Bronstein;S. Zafeiriou

文献摘要

相似文献

标准的配准算法需要在仔细的预处理和手动调整后,独立应用于每个表面进行配准。最近,出现了基于学习的方法,将新扫描的配准减少到使用先前训练的模型进行推理。潜在的好处是多方面的:推理通常比解决一个困难的优化问题的新实例快几个数量级,深度学习模型可以对噪声和腐败具有鲁棒性,并且经过训练的模型可以重新用于其他任务,例如通过迁移学习。在本文中,我们将配准任务转换为一个表面到表面的翻译问题,并设计了一个模型,以可靠地捕获潜在的几何信息直接从原始的3D人脸扫描。我们引入了Shape-My-Face(SMF),这是一种功能强大的编码器-解码器架构,基于改进的点云编码器,一种新颖的视觉注意力机制,具有跳过连接的图形卷积解码器,以及一种专门的嘴模型,我们可以将其与网格卷积平滑地集成在一起。与之前用于面部扫描的非刚性配准的最先进的学习算法相比,SMF只需要将原始数据与预定义的面部模板刚性对齐(缩放)。此外,我们的模型提供了具有最小监督的拓扑合理的网格,提供了更快的训练时间,具有数量级更少的可训练参数,对噪声更具鲁棒性,并且可以推广到以前看不见的数据集。我们根据不同的数据对注册质量进行广泛评估。我们证明了我们的模型的鲁棒性和通用性,在野外的人脸扫描在不同的方式,传感器类型和分辨率。最后,我们表明,通过学习注册扫描,SMF产生一个混合的线性和非线性变形模型。对SMF潜在空间的操纵允许形状生成和变形应用,例如野生表达转移。我们在人脸数据集上训练SMF,该数据集包括9个大型商品硬件数据库。
Standard registration algorithms need to be independently applied to each surface to register, following careful pre-processing and hand-tuning. Recently, learning-based approaches have emerged that reduce the registration of new scans to running inference with a previously-trained model. The potential benefits are multifold: inference is typically orders of magnitude faster than solving a new instance of a difficult optimization problem, deep learning models can be made robust to noise and corruption, and the trained model may be re-used for other tasks, e.g. through transfer learning. In this paper, we cast the registration task as a surface-to-surface translation problem, and design a model to reliably capture the latent geometric information directly from raw 3D face scans. We introduce Shape-My-Face (SMF), a powerful encoder-decoder architecture based on an improved point cloud encoder, a novel visual attention mechanism, graph convolutional decoders with skip connections, and a specialized mouth model that we smoothly integrate with the mesh convolutions. Compared to the previous state-of-the-art learning algorithms for non-rigid registration of face scans, SMF only requires the raw data to be rigidly aligned (with scaling) with a pre-defined face template. Additionally, our model provides topologically-sound meshes with minimal supervision, offers faster training time, has orders of magnitude fewer trainable parameters, is more robust to noise, and can generalize to previously unseen datasets. We extensively evaluate the quality of our registrations on diverse data. We demonstrate the robustness and generalizability of our model with in-the-wild face scans across different modalities, sensor types, and resolutions. Finally, we show that, by learning to register scans, SMF produces a hybrid linear and non-linear morphable model. Manipulation of the latent space of SMF allows for shape generation, and morphing applications such as expression transfer in-the-wild. We train SMF on a dataset of human faces comprising 9 large-scale databases on commodity hardware.