Emotion Recognition With Audio, Video, EEG, and EMG: A Dataset and Baseline Approaches

Emotion Recognition With Audio, Video, EEG, and EMG: A Dataset and Baseline Approaches
复制标题

DOI:
10.1109/access.2022.3146729
复制
发表时间:
2022-01-01
期刊:
影响因子:
3.9
通讯作者:
Zhu, Zhigang
Zhu, Zhigang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chen, Jin;Ro, Tony;Zhu, Zhigang

文献摘要

被引文献

相似文献

本文描述了一个新的提出的多模态情感数据集,并比较人类情感分类的基础上四种不同的方式-音频,视频,肌电图(EMG),脑电图(EEG)。报告的结果与几个基线的方法,使用各种特征提取技术和机器学习算法。首先,我们收集了来自11名人类受试者的数据集,这些受试者表达了六种基本情绪和一种中性情绪。然后,我们使用主成分分析,自动编码器,卷积网络和梅尔频率倒谱系数(MFCC),一些独特的个别模态提取特征。许多基线模型已被应用于比较情感识别中的分类性能,包括k-最近邻(KNN),支持向量机(SVM),随机森林,多层感知器(MLP),长短期记忆(LSTM)模型和卷积神经网络(CNN)。我们的结果表明,自举生物传感器信号(即,EMG和EEG)可以通过减少噪声来大大提高情感分类性能。相比之下,传统的KNN获得了最好的分类结果,而使用LSTM可以更好地分类人类情感的音频和图像序列。
This paper describes a new posed multimodal emotional dataset and compares human emotion classification based on four different modalities - audio, video, electromyography (EMG), and electroencephalography (EEG). The results are reported with several baseline approaches using various feature extraction techniques and machine-learning algorithms. First, we collected a dataset from 11 human subjects expressing six basic emotions and one neutral emotion. We then extracted features from each modality using principal component analysis, autoencoder, convolution network, and mel-frequency cepstral coefficient (MFCC), some unique to individual modalities. A number of baseline models have been applied to compare the classification performance in emotion recognition, including k-nearest neighbors (KNN), support vector machines (SVM), random forest, multilayer perceptron (MLP), long short-term memory (LSTM) model, and convolutional neural network (CNN). Our results show that bootstrapping the biosensor signals (i.e., EMG and EEG) can greatly increase emotion classification performance by reducing noise. In contrast, the best classification results were obtained by a traditional KNN, whereas audio and image sequences of human emotions could be better classified using LSTM.