SAT-Net: Self-Attention and Temporal Fusion for Facial Action Unit Detection

SAT-Net: Self-Attention and Temporal Fusion for Facial Action Unit Detection
复制标题

DOI:
10.1109/icpr48806.2021.9413260
复制
发表时间:
2021-01
期刊:
2020 25th International Conference on Pattern Recognition (ICPR)
影响因子:
--
通讯作者:
Zhihua Li;Zheng Zhang;L. Yin
Zhihua Li;Zheng Zhang;L. Yin
中科院分区:
其他
文献类型:
--
作者:
Zhihua Li;Zheng Zhang;L. Yin

文献摘要

相似文献

近年来,基于深度空间学习模型的人脸动作单元检测研究取得了令人瞩目的成绩,但由于缺乏对AUS跨时间时间信息的利用,其学习能力远未达到完全发挥的地步。由于一帧中AU的出现很可能与时间序列中的先前帧相关,因此探索各帧中AU的时间相关性成为本工作的关键动机。在本文中,我们提出了一种新的时间融合和AU监督的自我注意网络(所谓的SAT网)来解决AU检测问题。首先,我们将序列的深层特征输入卷积LSTM网络,并将先前的时间信息融合到最后一帧的特征映射中,继续学习AU的发生。其次,考虑到AU检测问题是一个多标签分类问题,即单个标签只依赖于特定的人脸区域,我们提出了一种新的自学习注意力模板,通过对每个AU的个体注意力掩码的学习,将每个AU的检测集中到人脸的部分区域,从而在不损失任何空间关系的情况下提高了AU的独立性。我们在两个基准数据库(BP4D和DISFA)上的大量实验表明,该框架在AU检测上取得了比目前最先进的结果更好的结果。
Research on facial action unit detection has shown remarkable performances by using deep spatial learning models in recent years, however, it is far from reaching its full capacity in learning due to the lack of use of temporal information of AUs across time. Since the AU occurrence in one frame is highly likely related to previous frames in a temporal sequence, exploring temporal correlation of AUs across frames becomes a key motivation of this work. In this paper, we propose a novel temporal fusion and AU-supervised self-attention network (a socalled SAT-Net) to address the AU detection problem. First of all, we input the deep features of a sequence into a convolutional LSTM network and fuse the previous temporal information into the feature map of the last frame, and continue to learn the AU occurrence. Second, considering the AU detection problem is a multi-label classification problem that individual label depends only on certain facial areas, we propose a new self-learned attention mask by focusing the detection of each AU on parts of facial areas through the learning of individual attention mask for each AU, thus increasing the AU independence without the loss of any spatial relations. Our extensive experiments show that the proposed framework achieves better results of AU detection over the state-of-the-arts on two benchmark databases (BP4D and DISFA).