Speech Emotion Recognition with Heterogeneous Feature Unification of Deep Neural Network

Speech Emotion Recognition with Heterogeneous Feature Unification of Deep Neural Network
复制标题

深度神经网络异构特征统一的语音情感识别

DOI:
10.3390/s19122730
复制
发表时间:
2019-06-02
期刊:
影响因子:
3.9
通讯作者:
Li, Chunguang
Li, Chunguang
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Jiang, Wei;Wang, Zheng;Li, Chunguang

文献摘要

被引文献

相似文献

语音情感识别是一项具有挑战性的任务,因为语音特征与人类情感之间存在着很大的差距,而语音情感识别在很大程度上依赖于为给定的识别任务提取的区别性声学特征。我们提出了一种新的深度神经结构来从异构声学特征组中提取信息特征表示,这些特征组可能包含冗余和不相关的信息,导致情绪识别性能低。在获得信息特征后,训练融合网络共同学习判别性声学特征表示,并使用支持向量机作为最终分类器完成识别任务。在IEMOCAP数据集上的实验结果表明,与现有的最先进的方法相比,所提出的架构提高了识别性能,达到了64%的准确率。
Automatic speech emotion recognition is a challenging task due to the gap between acoustic features and human emotions, which rely strongly on the discriminative acoustic features extracted for a given recognition task. We propose a novel deep neural architecture to extract the informative feature representations from the heterogeneous acoustic feature groups which may contain redundant and unrelated information leading to low emotion recognition performance in this work. After obtaining the informative features, a fusion network is trained to jointly learn the discriminative acoustic feature representation and a Support Vector Machine (SVM) is used as the final classifier for recognition task. Experimental results on the IEMOCAP dataset demonstrate that the proposed architecture improved the recognition performance, achieving accuracy of 64% compared to existing state-of-the-art approaches.