Convolutional aggregation of local evidence for large pose face alignment

Convolutional aggregation of local evidence for large pose face alignment
复制标题

DOI:
10.5244/c.30.86
复制
发表时间:
2016-09
期刊:
--
影响因子:
--
通讯作者:
Adrian Bulat;Yorgos Tzimiropoulos
Adrian Bulat;Yorgos Tzimiropoulos
中科院分区:
其他
文献类型:
--
作者:
Adrian Bulat;Yorgos Tzimiropoulos

文献摘要

被引文献

相似文献

无约束的人脸对齐方法必须满足两个要求:它们必须不依赖于精确的初始化/人脸检测,并且它们应该在整个面部姿态谱上表现良好。据我们所知,目前还没有一种方法能很好地满足这些要求,在本文中,我们提出了局部证据的卷积聚合(CALE),这是一种专门为解决这两个问题而设计的卷积神经网络(CNN)架构。特别是,为了消除对精确人脸检测的要求,我们的系统首先进行面部部分检测,为每个面部标志(局部证据)的位置提供置信度分数。接下来,我们的系统通过联合回归将这些分数地图与早期CNN特征聚合在一起,以优化地标的位置。除了扮演图形模型的角色,CNN回归是我们系统的一个关键特征,引导网络依赖上下文来预测遮挡的地标的位置,通常在非常大的姿势中遇到。整个系统在中间监督下进行端到端训练。当应用于AFLW-PIFA(迄今为止最具挑战性的人脸对齐测试集)时,与其他最近发表的大姿态人脸对齐方法相比,我们的方法提供了超过50%的定位精度增益。除了人脸,我们还证明了CALE在处理形状和外观的巨大变化方面是有效的,这通常是在动物的脸上遇到的。
Methods for unconstrained face alignment must satisfy two requirements: they must not rely on accurate initialisation/face detection and they should perform equally well for the whole spectrum of facial poses. To the best of our knowledge, there are no methods meeting these requirements to satisfactory extent, and in this paper, we propose Convolutional Aggregation of Local Evidence (CALE), a Convolutional Neural Network (CNN) architecture particularly designed for addressing both of them. In particular, to remove the requirement for accurate face detection, our system firstly performs facial part detection, providing confidence scores for the location of each of the facial landmarks (local evidence). Next, these score maps along with early CNN features are aggregated by our system through joint regression in order to refine the landmarks’ location. Besides playing the role of a graphical model, CNN regression is a key feature of our system, guiding the network to rely on context for predicting the location of occluded landmarks, typically encountered in very large poses. The whole system is trained end-to-end with intermediate supervision. When applied to AFLW-PIFA, the most challenging human face alignment test set to date, our method provides more than 50% gain in localisation accuracy when compared to other recently published methods for large pose face alignment. Going beyond human faces, we also demonstrate that CALE is effective in dealing with very large changes in shape and appearance, typically encountered in animal faces.