Robust semantic analysis by synthesis of 3D facial motion

Robust semantic analysis by synthesis of 3D facial motion
复制标题

DOI:
10.1109/fg.2011.5771336
复制
发表时间:
2011-03
期刊:
Face and Gesture 2011
影响因子:
--
通讯作者:
M. Breidt;H. Bülthoff;Cristóbal Curio
M. Breidt;H. Bülthoff;Cristóbal Curio
中科院分区:
其他
文献类型:
--
作者:
M. Breidt;H. Bülthoff;Cristóbal Curio

文献摘要

被引文献

相似文献

丰富的人脸模型已经对计算机视觉、感知研究以及计算机图形学和动画等领域产生了巨大的影响。诸如连续性、语义和直观控制之类的属性是理想的属性,但很难实现。为了实现构建这种高质量人脸模型的目标,我们提出了一种基于3D模型的合成分析方法,该方法能够对3D面部表面进行参数化,并且可以估计语义上有意义的组件的状态,即使是从嘈杂的深度数据中,例如由飞行时间(ToF)相机或Microsoft Kinect等设备产生的深度数据。在核心,我们提出了一个专门的三维变形模型(3DMM)的面部表情分析和合成。与许多其他模型相比,我们的模型来自于一个大型的局部面部变形语料库,这些变形被记录为来自多个身份的3D扫描。这使我们能够分析非结构化的动态3D扫描数据,使用修改的迭代最近点模型拟合过程,其次是约束动作单元模型回归,从而产生语义上有意义的面部变形时间过程。我们展示了我们的3DMMs的生成能力,从ToF相机的高质量和低质量的表面数据的面部表面重建。同时记录的面部运动,使用被动立体声和嘈杂的飞行时间相机的分析表明恢复的面部语义的良好协议。
Rich face models already have a large impact on the fields of computer vision, perception research, as well as computer graphics and animation. Attributes such as descriptiveness, semantics, and intuitive control are desirable properties but hard to achieve. Towards the goal of building such high-quality face models, we present a 3D model-based analysis-by-synthesis approach that is able to parameterize 3D facial surfaces, and that can estimate the state of semantically meaningful components, even from noisy depth data such as that produced by Time-of-Flight (ToF) cameras or devices such as Microsoft Kinect. At the core, we present a specialized 3D morphable model (3DMM) for facial expression analysis and synthesis. In contrast to many other models, our model is derived from a large corpus of localized facial deformations that were recorded as 3D scans from multiple identities. This allows us to analyze unstructured dynamic 3D scan data using a modified Iterative Closest Point model fitting process, followed by a constrained Action Unit model regression, resulting in semantically meaningful facial deformation time courses. We demonstrate the generative capabilities of our 3DMMs for facial surface reconstruction on high and low quality surface data from a ToF camera. The analysis of simultaneous recordings of facial motion using passive stereo and noisy Time-of-Flight camera shows good agreement of the recovered facial semantics.