Multimodal Neurophysiological Transformer for Emotion Recognition

Multimodal Neurophysiological Transformer for Emotion Recognition
复制标题

DOI:
10.1109/embc48229.2022.9871421
复制
发表时间:
2022-07
期刊:
2022 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC)
影响因子:
--
通讯作者:
Sharath C. Koorathota;Zain Khan;Pawan Lapborisuth;P. Sajda
Sharath C. Koorathota;Zain Khan;Pawan Lapborisuth;P. Sajda
中科院分区:
其他
文献类型:
--
作者:
Sharath C. Koorathota;Zain Khan;Pawan Lapborisuth;P. Sajda

文献摘要

相似文献

了解神经功能通常需要多种形式的数据,包括电生理数据,成像技术和人口统计调查。在本文中,我们介绍了一种新的神经生理模型,以解决建模多模态数据的主要挑战。首先,我们通过解决可变采样率的问题来避免原始信号和提取的频域特征之间的不对准问题。其次,我们通过与其他模态的“交叉注意”来编码模态。最后,我们利用我们的父Transformer架构的属性来建模跨模态的段之间的远程依赖关系,并评估中间权重,以更好地理解源信号如何影响预测。我们应用我们的多模态神经生理学Transformer(MNT)来预测现有开源数据集中的效价和唤醒。在非对齐多模态时间序列上的实验表明,我们的模型在分类任务中表现相似,并且在某些情况下优于现有方法。此外,定性分析表明,MNT是能够模拟神经的自主活动在预测唤醒的影响。我们的架构有可能针对各种下游任务进行微调,包括BCI系统。
Understanding neural function often requires multiple modalities of data, including electrophysiogical data, imaging techniques, and demographic surveys. In this paper, we introduce a novel neurophysiological model to tackle major challenges in modeling multimodal data. First, we avoid non-alignment issues between raw signals and extracted, frequency-domain features by addressing the issue of variable sampling rates. Second, we encode modalities through “cross-attention” with other modalities. Lastly, we utilize properties of our parent transformer architecture to model long-range dependencies between segments across modalities and assess intermediary weights to better understand how source signals affect prediction. We apply our Multimodal Neurophysiological Transformer (MNT) to predict valence and arousal in an existing open-source dataset. Experiments on non-aligned multimodal time-series show that our model performs similarly and, in some cases, outperforms existing methods in classification tasks. In addition, qualitative analysis suggests that MNT is able to model neural influences on autonomic activity in predicting arousal. Our architecture has the potential to be fine-tuned to a variety of downstream tasks, including for BCI systems.