Adversarial Learning Semantic Volume for 2D/3D Face Shape Regression in the Wild

Adversarial Learning Semantic Volume for 2D/3D Face Shape Regression in the Wild
复制标题

DOI:
10.1109/tip.2019.2911114
复制
发表时间:
2019-04
影响因子:
10.6
通讯作者:
Hongwen Zhang;Qi Li;Zhenan Sun
Hongwen Zhang;Qi Li;Zhenan Sun
中科院分区:
计算机科学1区
文献类型:
--
作者:
Hongwen Zhang;Qi Li;Zhenan Sun

文献摘要

相似文献

基于回归的方法已经彻底改变了2D地标定位,利用了深度神经网络和野外的大量注释数据集。然而,它仍然具有挑战性的3D地标定位由于缺乏注释的数据集和模糊的性质下的3D视角的地标。本文回顾了基于回归的方法,并提出了一个对抗体素和坐标回归框架,用于现实世界场景中的2D和3D面部标志定位。首先,引入语义体积表示来编码作为3D地标的位置的每体素可能性。然后,设计一个端到端的流水线来联合回归所提出的体积表示和坐标向量。这样的管道不仅增强了预测的鲁棒性和准确性,而且还统一了2D和3D地标定位,以便可以同时利用2D和3D数据集。此外,利用对抗学习策略将从合成数据集学习到的3D结构提取到弱监督设置下的真实世界数据集,其中提出了辅助回归模型来鼓励网络为合成和真实世界图像产生合理的预测。我们的方法的有效性进行了验证基准数据集3DFAW和AFLW 2000 -3D的二维和三维面部标志定位任务。实验结果表明,该方法取得了显着的改善,比以前的国家的最先进的方法。
Regression-based methods have revolutionized 2D landmark localization with the exploitation of deep neural networks and massive annotated datasets in the wild. However, it remains challenging for 3D landmark localization due to the lack of annotated datasets and the ambiguous nature of landmarks under the 3D perspective. This paper revisits regression-based methods and proposes an adversarial voxel and coordinate regression framework for 2D and 3D facial landmark localization in real-world scenarios. First, a semantic volumetric representation is introduced to encode the per-voxel likelihood of positions being the 3D landmarks. Then, an end-to-end pipeline is designed to jointly regress the proposed volumetric representation and the coordinate vector. Such a pipeline not only enhances the robustness and accuracy of the predictions but also unifies the 2D and 3D landmark localization so that the 2D and 3D datasets could be utilized simultaneously. Further, an adversarial learning strategy is exploited to distill 3D structure learned from synthetic datasets to real-world datasets under weakly supervised settings, where an auxiliary regression discriminator is proposed to encourage the network to produce plausible predictions for both the synthetic and real-world images. The effectiveness of our method is validated on benchmark datasets 3DFAW and AFLW2000-3D for both 2D and 3D facial landmark localization tasks. The experimental results show that the proposed method achieves significant improvements over the previous state-of-the-art methods.