Video modeling and learning on Riemannian manifold for emotion recognition in the wild

Video modeling and learning on Riemannian manifold for emotion recognition in the wild
复制标题

DOI:
10.1007/s12193-015-0204-5
复制
发表时间:
2015-11
影响因子:
2.9
通讯作者:
Mengyi Liu;Ruiping Wang;Shaoxin Li;Zhiwu Huang;S. Shan;Xilin Chen
Mengyi Liu;Ruiping Wang;Shaoxin Li;Zhiwu Huang;S. Shan;Xilin Chen
中科院分区:
计算机科学3区
文献类型:
--
作者:
Mengyi Liu;Ruiping Wang;Shaoxin Li;Zhiwu Huang;S. Shan;Xilin Chen

文献摘要

相似文献

在本文中,我们提出了在野外挑战中接受情感识别的方法(EMotiW)。面临的挑战是在真实环境下对视频片段中的人类主体行为的情感进行自动分类。在我们的方法中,每个视频片段可以分别由三种类型的图像集模型(即线性子空间、协方差矩阵和高斯分布)来表示,这些图像集模型都可以看作是驻留在一些黎曼流形上的点。然后对这些集合模型分别使用不同的黎曼核进行相似性/距离度量。在分类方面,对核支持向量机、Logistic回归和偏最小二乘三种分类器进行了比较。最后,在决策层对不同核函数和不同模式(视频和音频)学习的分类器进行最优融合,以进一步提高性能。我们对EMotiW 2014挑战数据(包括验证集和盲测试集)进行了广泛的评估,并评估了我们正在进行的不同组件的影响。据观察,我们的方法达到了迄今为止报道的最好的性能。为了进一步评估泛化能力,我们还在EMotiW 2013数据和两个著名的实验室控制数据库CK+和MMI上进行了实验。结果表明,该框架的性能明显优于现有的方法。
In this paper, we present the method for our submission to the emotion recognition in the wild challenge (EmotiW). The challenge is to automatically classify the emotions acted by human subjects in video clips under real-world environment. In our method, each video clip can be represented by three types of image set models (i.e. linear subspace, covariance matrix, and Gaussian distribution) respectively, which can all be viewed as points residing on some Riemannian manifolds. Then different Riemannian kernels are employed on these set models correspondingly for similarity/distance measurement. For classification, three types of classifiers, i.e. kernel SVM, logistic regression, and partial least squares, are investigated for comparisons. Finally, an optimal fusion of classifiers learned from different kernels and different modalities (video and audio) is conducted at the decision level for further boosting the performance. We perform extensive evaluations on the EmotiW 2014 challenge data (including validation set and blind test set), and evaluate the effects of different components in our pipeline. It is observed that our method has achieved the best performance reported so far. To further evaluate the generalization ability, we also perform experiments on the EmotiW 2013 data and two well-known lab-controlled databases: CK+ and MMI. The results show that the proposed framework significantly outperforms the state-of-the-art methods.