FACS3D-Net: 3D Convolution based Spatiotemporal Representation for Action Unit Detection

FACS3D-Net: 3D Convolution based Spatiotemporal Representation for Action Unit Detection
复制标题

DOI:
10.1109/acii.2019.8925514
复制
发表时间:
2019-09
期刊:
2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII)
影响因子:
--
通讯作者:
Le Yang;Itir Onal Ertugrul;J. Cohn;Z. Hammal;D. Jiang;H. Sahli
Le Yang;Itir Onal Ertugrul;J. Cohn;Z. Hammal;D. Jiang;H. Sahli
中科院分区:
其他
文献类型:
--
作者:
Le Yang;Itir Onal Ertugrul;J. Cohn;Z. Hammal;D. Jiang;H. Sahli

文献摘要

相似文献

大多数自动面部动作单元(Au)检测方法只考虑空间信息,而忽略了Au动态。对于人类来说,动态改善了Au感知。算法也是如此吗?为了利用Au动力学,最近在自动化Au检测方面的工作提出了一种顺序时空方法:使用2D CNN建模空间信息,然后使用LSTM(长短期记忆)建模时间信息。受人类FACS编码器经验的启发,我们假设同时结合空间和时间信息将产生更强大的Au检测。为了实现这一目标,我们提出了同时集成3D和2D CNN的FACS 3D-Net。评估是在200名参与者的扩展BP 4D+数据库上进行的。FACS 3D-Net优于2D CNN和2D CNN-LSTM方法。可视化的学习表示表明,FACS 3D-Net与人类FACS编码器所关注的时空动态一致。据我们所知,这是第一个将3D CNN应用于Au检测问题的工作。
Most approaches to automatic facial action unit (AU) detection consider only spatial information and ignore AU dynamics. For humans, dynamics improves AU perception. Is same true for algorithms? To make use of AU dynamics, recent work in automated AU detection has proposed a sequential spatiotemporal approach: Model spatial information using a 2D CNN and then model temporal information using LSTM (Long-Short-Term Memory). Inspired by the experience of human FACS coders, we hypothesized that combining spatial and temporal information simultaneously would yield more powerful AU detection. To achieve this, we propose FACS3D-Net that simultaneously integrates 3D and 2D CNN. Evaluation was on the Expanded BP4D+ database of 200 participants. FACS3D-Net outperformed both 2D CNN and 2D CNN-LSTM approaches. Visualizations of learnt representations suggest that FACS3D-Net is consistent with the spatiotemporal dynamics attended to by human FACS coders. To the best of our knowledge, this is the first work to apply 3D CNN to the problem of AU detection.