Musical Emotion Recognition with Spectral Feature Extraction Based on a Sinusoidal Model with Model-Based and Deep-Learning Approaches

Musical Emotion Recognition with Spectral Feature Extraction Based on a Sinusoidal Model with Model-Based and Deep-Learning Approaches
复制标题

DOI:
10.3390/app10030902
复制
发表时间:
2020-02-01
影响因子:
2.7
通讯作者:
Park, Chung Hyuk
Park, Chung Hyuk
中科院分区:
综合性期刊4区
文献类型:
--
作者:
Xie, Baijun;Kim, Jonathan C.;Park, Chung Hyuk

文献摘要

被引文献

相似文献

提出了一种基于正弦模型的新光谱特征提取方法。该方法的重点是利用频率子带中的频谱峰值来表征音频信号的频谱形状。对提取的特征进行评估,以预测情绪维度的水平,即唤醒和效价。主成分回归、偏最小二乘回归和深度卷积神经网络(CNN)模型被用作情绪维度水平的预测模型。实验结果表明,提出的特征包含了普通基线特征可能不包含的额外光谱信息。由于音频信号的质量,特别是音色,在影响音乐中情感配价的感知方面起着重要作用,因此包含所提出的特征将有助于降低预测错误率。
This paper presents a method for extracting novel spectral features based on a sinusoidal model. The method is focused on characterizing the spectral shapes of audio signals using spectral peaks in frequency sub-bands. The extracted features are evaluated for predicting the levels of emotional dimensions, namely arousal and valence. Principal component regression, partial least squares regression, and deep convolutional neural network (CNN) models are used as prediction models for the levels of the emotional dimensions. The experimental results indicate that the proposed features include additional spectral information that common baseline features may not include. Since the quality of audio signals, especially timbre, plays a major role in affecting the perception of emotional valence in music, the inclusion of the presented features will contribute to decreasing the prediction error rate.