Deep Temporal-Spatial Aggregation for Video-Based Facial Expression Recognition

Deep Temporal-Spatial Aggregation for Video-Based Facial Expression Recognition
复制标题

DOI:
10.3390/sym11010052
复制
发表时间:
2019-01
期刊:
Symmetry
影响因子:
--
通讯作者:
Xianzhang Pan;Wenping Guo;Xiaoying Guo;Wenshu Li;Junjie Xu;Jinzhao Wu
Xianzhang Pan;Wenping Guo;Xiaoying Guo;Wenshu Li;Junjie Xu;Jinzhao Wu
中科院分区:
其他
文献类型:
--
作者:
Xianzhang Pan;Wenping Guo;Xiaoying Guo;Wenshu Li;Junjie Xu;Jinzhao Wu

文献摘要

相似文献

所提出的方法具有30个流,即,15个空间流和15个时间流。每个空间流对应于每个时间流。因此,这项工作与对称性概念相关。由于视觉描述符和情感之间存在差距,对基于视频的人脸表情进行分类是一项困难的任务。为了弥补这一差距,提出了一种新的用于面部表情识别的视频描述符,以在整个视频范围内聚合空间和时间卷积特征。所设计的框架集成了一个国家的最先进的30流,并具有可训练的时空特征聚合层。该框架是端到端可训练的基于视频的面部表情识别。因此,该框架可以有效地避免对有限的情感视频数据集的过拟合,并且可训练策略可以学习更好地表示整个视频。研究了不同的时空特征池模式,并利用所提出的方法对时空流进行了最佳聚合。在两个公共数据库BAUM-1 s和eNTERFACE 05上的实验表明,该框架具有良好的性能,优于现有的策略。
The proposed method has 30 streams, i.e., 15 spatial streams and 15 temporal streams. Each spatial stream corresponds to each temporal stream. Therefore, this work correlates with the symmetry concept. It is a difficult task to classify video-based facial expression owing to the gap between the visual descriptors and the emotions. In order to bridge the gap, a new video descriptor for facial expression recognition is presented to aggregate spatial and temporal convolutional features across the entire extent of a video. The designed framework integrates a state-of-the-art 30 stream and has a trainable spatial–temporal feature aggregation layer. This framework is end-to-end trainable for video-based facial expression recognition. Thus, this framework can effectively avoid overfitting to the limited emotional video datasets, and the trainable strategy can learn to better represent an entire video. The different schemas for pooling spatial–temporal features are investigated, and the spatial and temporal streams are best aggregated by utilizing the proposed method. The extensive experiments on two public databases, BAUM-1s and eNTERFACE05, show that this framework has promising performance and outperforms the state-of-the-art strategies.