Real-time expression transfer for facial reenactment

Real-time expression transfer for facial reenactment
复制标题

DOI:
10.1145/2816795.2818056
复制
发表时间:
2015-10
期刊:
ACM Transactions on Graphics (TOG)
影响因子:
--
通讯作者:
Justus Thies;M. Zollhöfer;M. Nießner;Levi Valgaerts;M. Stamminger;C. Theobalt
Justus Thies;M. Zollhöfer;M. Nießner;Levi Valgaerts;M. Stamminger;C. Theobalt
中科院分区:
其他
文献类型:
--
作者:
Justus Thies;M. Zollhöfer;M. Nießner;Levi Valgaerts;M. Stamminger;C. Theobalt

文献摘要

被引文献

相似文献

我们提出了一种方法,用于从源视频中的演员到目标视频中的演员的面部表情的实时传输,从而使目标演员的面部表情的ad-hoc控制。我们的方法的新奇在于转移和逼真的重新渲染到目标视频的面部变形和细节的方式,新合成的表情几乎无法区分从一个真实的视频。为了实现这一目标,我们使用商品RGB-D传感器实时准确地捕获源和目标对象的面部表现。对于每一帧,我们共同拟合的身份,表情和皮肤反射的参数模型的输入颜色和深度数据,并重建场景照明。对于表达式转换,我们在参数空间中计算源表达式和目标表达式之间的差异,并修改目标参数以匹配源表达式。一个主要的挑战是令人信服的重新渲染合成的目标脸到相应的视频流。这需要仔细考虑照明和阴影设计,两者都必须符合现实世界的环境。我们在现场设置中演示了我们的方法,在现场设置中,我们修改了视频会议馈送,使得不同人的面部表情(例如,翻译器)实时匹配。
We present a method for the real-time transfer of facial expressions from an actor in a source video to an actor in a target video, thus enabling the ad-hoc control of the facial expressions of the target actor. The novelty of our approach lies in the transfer and photorealistic re-rendering of facial deformations and detail into the target video in a way that the newly-synthesized expressions are virtually indistinguishable from a real video. To achieve this, we accurately capture the facial performances of the source and target subjects in real-time using a commodity RGB-D sensor. For each frame, we jointly fit a parametric model for identity, expression, and skin reflectance to the input color and depth data, and also reconstruct the scene lighting. For expression transfer, we compute the difference between the source and target expressions in parameter space, and modify the target parameters to match the source expressions. A major challenge is the convincing re-rendering of the synthesized target face into the corresponding video stream. This requires a careful consideration of the lighting and shading design, which both must correspond to the real-world environment. We demonstrate our method in a live setup, where we modify a video conference feed such that the facial expressions of a different person (e.g., translator) are matched in real-time.