Cross-individual affective detection using EEG signals with audio-visual embedding

Cross-individual affective detection using EEG signals with audio-visual embedding
复制标题

使用带有视听嵌入的脑电图信号进行跨个体情感检测

DOI:
10.1016/j.neucom.2022.09.078
复制
发表时间:
2022
期刊:
影响因子:
6
通讯作者:
Zhiguo Zhang
Zhiguo Zhang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhen Liang;Xihao Zhang;Rushuang Zhou;Li Zhang;Linling Li;Gan Huang;Zhiguo Zhang

文献摘要

相似文献

情感计算是一个新兴的跨学科研究领域,它为识别、理解和表达人类情感提供了巨大的潜力。近年来,多模态分析在情感研究中越来越受欢迎,它可以基于来自不同数据模态的多样性和互补性信息,提供更全面的情感动态视图。然而,目前的多模态分析方法的稳定性和可泛化性还没有得到充分的发展。本文提出了一种新的跨个体情感检测的多模态分析方法(EEG- ave: EEG with视听嵌入),该方法利用EEG信号识别与情感相关的个体偏好,利用视听信息估计多媒体内容中涉及的内在情感。EEG-AVE主要由两个模块组成。在基于脑电的个体偏好预测模块中,开发了一种多尺度域对抗神经网络来探索个体间共享的动态、信息丰富、域不变的脑电特征。在基于视频的内在情感估计模块中,提出了一种基于深度视听特征的超图聚类方法来检测语义视听特征与情感之间的潜在关系。通过嵌入模型,估计的个人偏好和内在情绪都被纳入共享权重,并进一步有助于个体之间的情感检测。在两个知名的情绪数据库上进行的实验表明,所提出的EEG-AVE模型在留下一个个体的交叉验证个体独立评估协议下取得了更好的性能。结果表明,EEG-AVE是一种有效的模型,具有良好的可靠性和泛化性,对情感计算中多模态分析的发展具有实际意义。
Affective computing is an increasing interdisciplinary research field that provides great potential to recognize, understand and express human emotions. Recently, multimodal analysis starts to gain more popularity in affective studies, which could provide a more comprehensive view of emotion dynamics based on the diverse and complementary information from different data modalities. However, the stability and generalizability of current multimodal analysis methods have not been thoroughly developed yet. In this paper, we propose a novel multimodal analysis method ( EEG-AVE : EEG with audio-visual embedding) for cross-individual affective detection, where EEG signals are exploited to identify the emotion-related individual preferences and audio-visual information is leveraged to estimate the intrinsic emotions involved in the multimedia content. EEG-AVE is composed of two main modules. For EEG-based individual preferences prediction module, a multi-scale domain adversarial neural network is developed to explore the shared dynamic, informative, and domain-invariant EEG features across individuals. For video-based intrinsic emotions estimation module, a deep audio-visual feature-based hypergraph clustering method is proposed to examine the latent relationship between semantic audio-visual features and emotions. Through an embedding model, both estimated individual preferences and intrinsic emotions are incorporated with shared weights and further contribute to affective detection across individuals. Experiments on two well-known emotional databases indicate that the proposed EEG-AVE model achieves a better performance under a leave-one-individual-out cross-validation individual-independent evaluation protocol. The results demonstrate that EEG-AVE is an effective model with good reliability and generalizability, which has practical significance in the development of multimodal analysis in affective computing.