Learning-by-Novel-View-Synthesis for Full-Face Appearance-Based 3D Gaze Estimation

Learning-by-Novel-View-Synthesis for Full-Face Appearance-Based 3D Gaze Estimation
复制标题

DOI:
10.1109/cvprw56347.2022.00546
复制
发表时间:
2022-01
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
Jiawei Qin;Takuru Shimoyama;Yusuke Sugano
Jiawei Qin;Takuru Shimoyama;Yusuke Sugano
中科院分区:
其他
文献类型:
--
作者:
Jiawei Qin;Takuru Shimoyama;Yusuke Sugano

文献摘要

相似文献

尽管最近在基于外观的注视估计技术方面取得了进展,但对覆盖目标头部姿势和注视分布的训练数据的需求仍然是实际部署的关键挑战。本文研究了一种基于单目三维人脸重建的视线估计训练数据合成方法。与使用多视图重建,照片般逼真的CG模型或生成神经网络的先前工作不同,我们的方法可以操纵和扩展现有训练数据的头部姿势范围,而无需任何额外的要求。我们引入了一个投影匹配过程,将重建的3D人脸网格与相机坐标系对齐,并合成具有准确凝视标签的人脸图像。我们还提出了一个面具引导的视线估计模型和数据增强策略,以进一步提高估计精度,利用合成训练数据。使用多个公共数据集的实验表明,我们的方法显着提高了具有挑战性的跨数据集设置与非重叠的凝视分布的估计性能。
Despite recent advances in appearance-based gaze estimation techniques, the need for training data that covers the target head pose and gaze distribution remains a crucial challenge for practical deployment. This work examines a novel approach for synthesizing gaze estimation training data based on monocular 3D face reconstruction. Unlike prior works using multi-view reconstruction, photo-realistic CG models, or generative neural networks, our approach can manipulate and extend the head pose range of existing training data without any additional requirements. We introduce a projective matching procedure to align the reconstructed 3D facial mesh with the camera coordinate system and synthesize face images with accurate gaze labels. We also propose a mask-guided gaze estimation model and data augmentation strategies to further improve the estimation accuracy by taking advantage of synthetic training data. Experiments using multiple public datasets show that our approach significantly improves the estimation performance on challenging cross-dataset settings with non-overlapping gaze distributions.