AuxFormer: Robust Approach to Audiovisual Emotion Recognition

AuxFormer: Robust Approach to Audiovisual Emotion Recognition
复制标题

DOI:
10.1109/icassp43922.2022.9747157
复制
发表时间:
2022-05
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Lucas Goncalves;C. Busso
Lucas Goncalves;C. Busso
中科院分区:
其他
文献类型:
--
作者:
Lucas Goncalves;C. Busso

文献摘要

相似文献

在视听情感识别中,一个具有挑战性的任务是实现能够利用和融合多模态信息的神经网络架构,同时暂时对齐模态,处理缺失模态,并从所有模态中捕获信息,而不会在训练过程中丢失信息。这些要求对于实现模型鲁棒性和提高情绪识别任务的准确性是非常重要的。执行多模态融合的最新方法是使用变压器架构来正确融合和对齐模态。本研究提出了AuxFormer框架,它以一种原则性的方式解决了上述挑战。AuxFormer将变压器框架与辅助网络相结合。它利用共享损耗,从单独嵌入的单模态网络中注入信息。添加到主网络的额外视听信息层保留了在训练过程中可能丢失的信息。结果表明,AuxFormer架构在CREMA-D语料库上的宏观和微观F1Scores分别达到71.3%和71.7%。对于MSP-IMPROV语料库,AuxFormer的宏观和微观f1得分分别为70.4%和76.5%。两个语料库的结果都明显优于强基线,表明我们的框架受益于辅助网络。我们还表明,在非理想条件下(例如,缺少模式),我们的架构能够在仅音频和仅视频的场景下保持强大的性能,受益于优化的训练策略。
A challenging task in audiovisual emotion recognition is to implement neural network architectures that can leverage and fuse multimodal information while temporally aligning modalities, handling missing modalities, and capturing information from all modalities without losing information during training. These requirements are important to achieve model robustness and to increase accuracy on the emotion recognition task. A recent approach to perform multimodal fusion is to use the transformer architecture to properly fuse and align the modalities. This study proposes the AuxFormer framework, which addresses in a principled way the aforementioned challenges. AuxFormer combines the transformer framework with auxiliary networks. It uses shared losses to infuse information from single-modality networks that are separately embedded. The extra layer of audiovisual information added to our main network retains information that would otherwise be lost during training. The results show that the AuxFormer architecture achieves macro and micro F1Scores of 71.3% and 71.7%, respectively, on the CREMA-D corpus. For the MSP-IMPROV corpus, AuxFormer achieves a macro and micro F1-Scores of 70.4% and 76.5%, respectively. The results for both corpora are significantly better than strong baselines, indicating that our framework benefits from auxiliary networks. We also show that under non-ideal conditions (e.g., missing modalities) our architecture is able to sustain strong performance under audio-only and video-only scenarios, benefiting from a optimized training strategy.