Self-supervised Deformation Modeling for Facial Expression Editing

Self-supervised Deformation Modeling for Facial Expression Editing
复制标题

DOI:
10.1109/fg47880.2020.00115
复制
发表时间:
2019-11
期刊:
2020 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020)
影响因子:
--
通讯作者:
ShahRukh Athar;Zhixin Shu;D. Samaras
ShahRukh Athar;Zhixin Shu;D. Samaras
中科院分区:
其他
文献类型:
--
作者:
ShahRukh Athar;Zhixin Shu;D. Samaras

文献摘要

相似文献

深度生成模型最近在逼真的面部图像合成和编辑方面展示了令人印象深刻的结果。现有的基于神经网络的方法通常仅依靠纹理生成来编辑表情,而很大程度上忽略了运动信息。然而,面部表情本质上是肌肉运动的结果。在这项工作中,我们提出了一种新颖的端到端网络,它将面部编辑任务分为两个步骤:“动作编辑”步骤和“纹理编辑”步骤。在“动作编辑”步骤中,我们通过图像变形明确地模拟面部运动,将图像扭曲成所需的表情。在“纹理编辑”步骤中,我们生成必要的纹理,例如牙齿和阴影效果,以获得逼真的结果。我们基于物理的任务解缠系统设计允许每个步骤学习一个重点任务,因此不需要生成纹理来产生运动幻觉。我们的系统以自我监督的方式进行训练,不需要地面实况变形注释。我们的方法使用动作单元[8]作为面部表情的表示,在定性和定量评估方面提高了最先进的面部表情编辑性能。1.
Deep generative models have recently demonstrated impressive results in photo-realistic facial image synthesis and editing. Existing neural network-based approaches usually only rely on texture generation to edit expressions and largely neglect the motion information. However, facial expressions are inherently the result of muscle movement. In this work, we propose a novel end-to-end network that disentangles the task of facial editing into two steps: a “motionediting” step and a “texture-editing” step. In the “motionediting” step, we explicitly model facial movement through an image deformation, warping the image into the desired expression. In the “texture-editing” step, we generate the necessary textures, such as teeth and shading effects, for a photorealistic result. Our physically-based task-disentanglement system design allows each step to learn a focused task, and thus need not generate texture to hallucinate motion. Our system is trained in a self-supervised manner, requiring no ground truth deformation annotation. Using Action Units [8] as the representation for facial expression, our method improves the state-of-the-art facial expression editing performance in both qualitative and quantitative evaluations.1.