Design of a Convolutional Neural Network for Speech Emotion Recognition
Design of a Convolutional Neural Network for Speech Emotion Recognition
复制标题
用于语音情感识别的卷积神经网络设计
DOI:
10.1109/ictc49870.2020.9289227
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Do Hyun Kim
中科院分区:
文献类型:
--
作者:
Kyong Hee Lee;Do Hyun Kim
Regarding speech emotion recognition (SER) using voice, recognition accuracy increases as more data are employed. In particular, in the case of deep learning, a large amount of data is essential. However, when using an existing data set, the size of the data set is limited, and the length of the data constituting the data set can be inconsistent. The data set used in this paper consists of audio files of utterances of various lengths. In this paper, one-dimensional data was extracted from speech files, and two-dimensional mel-spectrogram images were extracted and trained using deep learning techniques such as a multi-layer perceptron (MLP) and a convolutional neural network (CNN). In addition, to improve the test accuracy, audio files were reduced to less than two seconds and preprocessed. Using the CNN, we obtained a test accuracy of approximately 60%.