Towards a dynamic expression recognition system under facial occlusion

Towards a dynamic expression recognition system under facial occlusion
复制标题

面向面部遮挡下的动态表情识别系统

DOI:
10.1016/j.patrec.2012.07.015
复制
发表时间:
2012-12-01
影响因子:
5.1
通讯作者:
Pietikainen, Matti
Pietikainen, Matti
中科院分区:
计算机科学3区
文献类型:
--
作者:
Huang, Xiaohua;Zhao, Guoying;Pietikainen, Matti

文献摘要

被引文献

相似文献

人脸遮挡是人脸表情识别中一个具有挑战性的研究课题。这导致需要开发一些有趣的面部表示和遮挡检测方法,以便将FER扩展到不受控制的环境。应该指出的是,以前的工作主要集中在这两个问题上,并在静态图像。因此,我们有动机提出一个完整的系统,包括面部表示,遮挡检测,并在视频序列中的多个特征融合。为了实现一个强大的面部表示,我们提出了一种方法,从眼睛,鼻子和嘴巴组件派生六个特征向量,形成一个面部表示。这些具有时间线索的特征由动态纹理和结构形状特征描述符生成。另一方面,遮挡检测仍然主要通过传统的分类器或模型比较来实现。最近,稀疏表示已被提出作为一种有效的方法,以防止遮挡,而它是与人脸识别在FER,除非使用适当的面部表示。因此,我们提出了一个评估表明,建议的面部表示是独立的面部身份。受Mercier等人(2007)的启发,我们将稀疏表示和残差统计用于图像序列的遮挡检测。针对将六个特征向量合并为一个特征向量会导致维数灾难的问题,提出了由融合模块和权值学习组成的多特征融合方法。在扩展的Cohn-Kanade数据库及其模拟数据库上的实验结果表明,我们的框架在正常视频中,特别是在部分遮挡视频中的FER性能优于最先进的方法。(c)2012爱思唯尔有限公司版权所有。
Facial occlusion is a challenging research topic in facial expression recognition (FER). This has resulted in the need to develop some interesting facial representations and occlusion detection methods in order to extend the FER to uncontrolled environments. It should be noted that most of the previous work focuses on these two issues separately, and on static images. We are thus motivated to propose a complete system consisting of facial representations, occlusion detection, and multiple feature fusion in video sequences. For achieving a robust facial representation, we propose an approach deriving six feature vectors from eyes, nose and mouth components to form a facial representation. These features with temporal cues are generated by the dynamic texture and structural shape feature descriptors. On the other hand, occlusion detection is still mainly realized by the traditional classifiers or model comparison. Recently, sparse representation has been proposed as an efficient method against occlusion, while it is correlated with facial identity in FER, unless using an appropriate facial representation. Thus, we present an evaluation demonstrating that the proposed facial representation is independent of facial identity. Inspired by Mercier et al. (2007), we then exploit the use of the sparse representation and residual statistics to occlusion detection of the image sequences. As concatenating six feature vectors into one causes the curse of dimensionality, we propose multiple feature fusion consisting of fusion module and weight learning. Experimental results on the Extended Cohn-Kanade database and its simulated database demonstrate that our framework outperforms the state-of-the-art methods for FER in normal videos, and especially, in partial occlusion videos. (c) 2012 Elsevier B.V. All rights reserved.