3-D Facial Expression Recognition via Attention-Based Multichannel Data Fusion Network

3-D Facial Expression Recognition via Attention-Based Multichannel Data Fusion Network
复制标题

通过基于注意力的多通道数据融合网络进行 3D 面部表情识别

DOI:
10.1109/tim.2021.3125972
复制
发表时间:
2021
影响因子:
5.6
通讯作者:
Fuji Ren
Fuji Ren
中科院分区:
工程技术2区
文献类型:
--
作者:
Yu Gu;Huan Yan;Xiang Zhang;Zhi Liu;Fuji Ren

文献摘要

相似文献

长期以来,面部表情一直被认为包含解读人类情感的有意义的非语言情感线索。最近,多模态 2-D + 3-D 融合方法由于其在各种空间通道中的细粒度面部描述而在面部表情识别(FER)中显示出巨大的潜力。然而,当前的工作主要依靠特征甚至分数级别的融合来寻找在不同渠道中传播的情感线索,并且可能由于缺乏焦点而错过关键信息。为此,我们提出了一种基于注意力的多通道数据融合网络(AMDFN),以更好地保存和发现此类关键的面部线索。更具体地说,我们首先将 3D 面部扫描映射到多通道图像,然后将它们融合到 ResNet18 主干中以获得分层情感特征。其次,我们利用层注意力模型来探索不同层特征之间的依赖关系,以学习有效情绪识别的辨别性情感线索。我们对两个广泛使用的数据集(即 Facescape 和 Bosphorus)进行的综合实验验证了我们的方法与几个最先进的竞争对手相比的性能。
Facial expression has long been recognized as containing meaningful nonverbal affective cues for decoding human emotions. Recently, multimodal 2-D + 3-D fusion method has shown significant potential in facial expression recognition (FER) due to its fine-grained face descriptions in various spatial channels. However, current work mainly relies on feature- or even score-level fusion to find emotion cues spread in different channels and may miss key information due to lack of focus. To this end, we propose an attention-based multichannel data fusion network (AMDFN) to better preserve and find such key facial cues. More specifically, we first map a 3-D face scan into multichannel images and then fuse them in a ResNet18 backbone to get layered emotion features. Second, we leverage a layer attention model to explore the dependencies between features of different layers to learn discriminative affective cues for effective emotion recognition. Our comprehensive experiments on two widely used datasets (i.e., Facescape and Bosphorus) have verified the performance of our approach compared to several state-of-the-art rivals.