Design of a Convolutional Neural Network for Speech Emotion Recognition

Design of a Convolutional Neural Network for Speech Emotion Recognition
复制标题

用于语音情感识别的卷积神经网络设计

DOI:
10.1109/ictc49870.2020.9289227
复制
发表时间:
2020
期刊:
2020 International Conference on Information and Communication Technology Convergence (ICTC)
影响因子:
--
通讯作者:
Do Hyun Kim
Do Hyun Kim
中科院分区:
--
文献类型:
--
作者:
Kyong Hee Lee;Do Hyun Kim

文献摘要

被引文献

相似文献

对于基于语音的语音情感识别(SER),随着数据量的增加,识别准确率也随之提高。特别是在深度学习的情况下,大量的数据是必不可少的。但是,当使用现有数据集时,数据集的大小是有限的,并且组成数据集的数据的长度可能不一致。本文使用的数据集由不同长度的语音文件组成。本文从语音文件中提取一维数据,并使用多层感知器(MLP)和卷积神经网络(CNN)等深度学习技术提取和训练二维梅尔谱图图像。此外,为了提高测试精度,音频文件被压缩到两秒以内并进行预处理。使用CNN,我们获得了大约60%的测试准确率。
Regarding speech emotion recognition (SER) using voice, recognition accuracy increases as more data are employed. In particular, in the case of deep learning, a large amount of data is essential. However, when using an existing data set, the size of the data set is limited, and the length of the data constituting the data set can be inconsistent. The data set used in this paper consists of audio files of utterances of various lengths. In this paper, one-dimensional data was extracted from speech files, and two-dimensional mel-spectrogram images were extracted and trained using deep learning techniques such as a multi-layer perceptron (MLP) and a convolutional neural network (CNN). In addition, to improve the test accuracy, audio files were reduced to less than two seconds and preprocessed. Using the CNN, we obtained a test accuracy of approximately 60%.