Peak-Piloted Deep Network for Facial Expression Recognition

Peak-Piloted Deep Network for Facial Expression Recognition
复制标题

DOI:
10.1007/978-3-319-46475-6_27
复制
发表时间:
2016-07
期刊:
影响因子:
16
通讯作者:
Xiangyu Zhao;Xiaodan Liang;Luoqi Liu;Teng Li;Yugang Han;N. Vasconcelos;Shuicheng Yan
Xiangyu Zhao;Xiaodan Liang;Luoqi Liu;Teng Li;Yugang Han;N. Vasconcelos;Shuicheng Yan
中科院分区:
生物学1区
文献类型:
--
作者:
Xiangyu Zhao;Xiaodan Liang;Luoqi Liu;Teng Li;Yugang Han;N. Vasconcelos;Shuicheng Yan

文献摘要

被引文献

相似文献

用于训练面部相关识别任务(如面部表情识别(FER))的深度网络的目标函数通常独立考虑每个样本。在这项工作中,我们提出了一种新的峰值引导深度网络(PPDN),它使用具有峰值表达的样本(简单样本)来监督相同类型和相同主题的非峰值表达样本(硬样本)的中间特征响应。因此,从非峰值表达到峰值表达的表达进化过程可以隐式地嵌入到网络中,以实现表达强度的不变性。一个特殊用途的反向传播过程,峰值梯度抑制(PGS),提出了网络训练。它将非峰值表达样本的中间层特征响应驱动到对应峰值表达样本的中间层特征响应,同时避免反向。这避免了由于非峰值表达对应物的干扰而降低峰值表达样本的识别能力。在Oulu-CASIA和CK+两个流行的FER数据集上进行了广泛的比较,证明了PPDN相对于最先进的FER方法的优越性,以及网络结构和优化策略的优势。此外,它表明,PPDN是一个通用的架构,可扩展到其他任务的峰值和非峰值样本的适当定义。这是验证的实验,显示国家的最先进的性能姿态不变的人脸识别,使用多PIE数据集。
Objective functions for training of deep networks for face-related recognition tasks, such as facial expression recognition (FER), usually consider each sample independently. In this work, we present a novel peak-piloted deep network (PPDN) that uses a sample with peak expression (easy sample) to supervise the intermediate feature responses for a sample of non-peak expression (hard sample) of the same type and from the same subject. The expression evolving process from non-peak expression to peak expression can thus be implicitly embedded in the network to achieve the invariance to expression intensities. A special-purpose back-propagation procedure, peak gradient suppression (PGS), is proposed for network training. It drives the intermediate-layer feature responses of non-peak expression samples towards those of the corresponding peak expression samples, while avoiding the inverse. This avoids degrading the recognition capability for samples of peak expression due to interference from their non-peak expression counterparts. Extensive comparisons on two popular FER datasets, Oulu-CASIA and CK+, demonstrate the superiority of the PPDN over state-of-the-art FER methods, as well as the advantages of both the network structure and the optimization strategy. Moreover, it is shown that PPDN is a general architecture, extensible to other tasks by proper definition of peak and non-peak samples. This is validated by experiments that show state-of-the-art performance on pose-invariant face recognition, using the Multi-PIE dataset.