Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation Fusion

Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation Fusion
复制标题

DOI:
10.1109/icassp39728.2021.9413608
复制
发表时间:
2021-06
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Jiaxing Liu;Sen Chen;Longbiao Wang;Zhilei Liu;Yahui Fu;Lili Guo;J. Dang
Jiaxing Liu;Sen Chen;Longbiao Wang;Zhilei Liu;Yahui Fu;Lili Guo;J. Dang
中科院分区:
其他
文献类型:
--
作者:
Jiaxing Liu;Sen Chen;Longbiao Wang;Zhilei Liu;Yahui Fu;Lili Guo;J. Dang

文献摘要

相似文献

由于音视频多通道情感识别(MER)比单通道情感识别具有更强的鲁棒性,因此受到了人们的广泛关注。表示融合算法的效率往往决定着MER算法的性能。虽然融合算法很多,但通常忽略了信息冗余和信息互补。提出了一种新的表示融合方法--胶囊图卷积网络(CapsGCN)。首先,在单峰表示学习后,提取的音频和视频表示分别用胶囊网络提取并封装到多模胶囊中。多式联运胶囊可以通过动态路由算法有效地减少数据冗余。其次,将多式联运胶囊及其相互关系和内部关系视为一种图结构。通过图卷积网络(GCN)学习图的结构,得到隐含的表示,为信息互补提供了一个很好的补充。最后,CapsGCN学习的多通道胶囊和隐藏关系表征被反馈给多头自我注意,以平衡来源表征和关系表征的贡献。为了验证CapsGCN的性能、表示的可视化、常用融合方法的结果和烧蚀研究,提供了所提出的CapsGCN。我们提出的融合方法在eNTERFACE05‘上达到了80.83%的准确率和80.23%的F1得分。
Due to the more robust characteristics compared to unimodal, audio-video multimodal emotion recognition (MER) has attracted a lot of attention. The efficiency of representation fusion algorithm often determines the performance of MER. Although there are many fusion algorithms, information redundancy and information complementarity are usually ignored. In this paper, we propose a novel representation fusion method, Capsule Graph Convolutional Network (CapsGCN). Firstly, after unimodal representation learning, the extracted audio and video representations are distilled by capsule network and encapsulated into multimodal capsules respectively. Multimodal capsules can effectively reduce data redundancy by the dynamic routing algorithm. Secondly, the multimodal capsules with their inter-relations and intra-relations are treated as a graph structure. The graph structure is learned by Graph Convolutional Network (GCN) to get hidden representation which is a good supplement for information complementarity. Finally, the multimodal capsules and hidden relational representation learned by CapsGCN are fed to multihead self-attention to balance the contributions of source representation and relational representation. To verify the performance, visualization of representation, the results of commonly used fusion methods, and ablation studies of the proposed CapsGCN are provided. Our proposed fusion method achieves 80.83% accuracy and 80.23% F1 score on eNTERFACE05’.