An Automatic Framework for Textured 3D Video-Based Facial Expression Recognition

An Automatic Framework for Textured 3D Video-Based Facial Expression Recognition
复制标题

DOI:
10.1109/taffc.2014.2330580
复制
发表时间:
2014-07-01
影响因子:
11.2
通讯作者:
Bennamoun, Mohammed
Bennamoun, Mohammed
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hayat, Munawar;Bennamoun, Mohammed

文献摘要

被引文献

相似文献

现有的大部分三维人脸表情识别研究都是使用静态三维网格进行的。人脸的3D视频被认为包含更多关于面部动态的信息,这对于表情识别非常关键。本文提出了一种全自动框架,利用纹理的3D视频的动态识别六个离散的面部表情。从训练视频的许多位置提取可变长度的局部视频补丁,并表示为格拉斯曼流形上的点。一个有效的基于图的谱聚类算法被用来分别聚类这些点的每一个表达式类。使用一个有效的格拉斯曼核函数,得到的聚类中心嵌入到再生核希尔伯特空间(RKHS),其中六个二进制SVM模型的学习。给定一个查询视频,我们从中提取视频补丁,将它们表示为流形上的点,并将这些点与学习的SVM模型进行匹配,然后采用基于投票的策略来决定查询视频的类别。所提出的框架也实现了并行的2D视频和分数级融合的2D和3D视频进行系统的性能改进。在BU 4DFE数据集上的实验结果表明,该系统对3D视频中的面部表情识别取得了非常高的分类准确率。
Most of the existing research on 3D facial expression recognition has been done using static 3D meshes. 3D videos of a face are believed to contain more information in terms of the facial dynamics which are very critical for expression recognition. This paper presents a fully automatic framework which exploits the dynamics of textured 3D videos for recognition of six discrete facial expressions. Local video-patches of variable lengths are extracted from numerous locations of the training videos and represented as points on the Grassmannian manifold. An efficient graph-based spectral clustering algorithm is used to separately cluster these points for every expression class. Using a valid Grassmannian kernel function, the resulting cluster centers are embedded into a Reproducing Kernel Hilbert Space (RKHS) where six binary SVM models are learnt. Given a query video, we extract video-patches from it, represent them as points on the manifold and match these points with the learnt SVM models followed by a voting based strategy to decide about the class of the query video. The proposed framework is also implemented in parallel on 2D videos and a score level fusion of 2D & 3D videos is performed for performance improvement of the system. The experimental results on BU4DFE data set show that the system achieves a very high classification accuracy for facial expression recognition from 3D videos.