Dynamic Speech Emotion Recognition with State-Space Models

Dynamic Speech Emotion Recognition with State-Space Models
复制标题

使用状态空间模型的动态语音情感识别

DOI:
10.1109/eusipco.2015.7362750
复制
发表时间:
2015
期刊:
Signal Processing Conference (EUSIPCO), 2015 23rd European
影响因子:
--
通讯作者:
and Gareth. W. Peters
and Gareth. W. Peters
中科院分区:
--
文献类型:
--
作者:
Konstantin Markov;Tomoko Matsui;Francois Septier;and Gareth. W. Peters

文献摘要

相似文献

从语音中自动识别情感主要集中在识别分类或静态的情感状态,但人类情感的频谱是连续的和时变的。本文提出了一种基于状态空间模型的动态语音情感识别系统。预测未知的情绪轨迹的影响空间跨越的唤醒,效价和优势(AV-D)描述符被铸造作为一个时间序列过滤任务。我们研究的状态空间模型包括一个标准的线性模型(卡尔曼滤波器),以及新的非线性,非参数高斯过程(GP)为基础的SSM。我们使用AVEC 2014数据库进行评估,该数据库提供了真实的A-V-D标签,允许分别学习状态和测量函数,从而简化模型训练。对于GP SSM的滤波,我们使用了两种近似方法:最近提出的解析方法和粒子滤波。所有的模型进行了评估,平均皮尔逊相关系数R和均方根误差(RMSE)。结果表明,使用相同的特征向量,GP SSM实现两倍高的相关性和两倍小的RMSE比卡尔曼滤波器。
Automatic emotion recognition from speech has been focused mainly on identifying categorical or static affect states, but the spectrum of human emotion is continuous and time-varying. In this paper, we present a recognition system for dynamic speech emotion based on state-space models (SSMs). The prediction of the unknown emotion trajectory in the affect space spanned by Arousal, Valence, and Dominance (A-V-D) descriptors is cast as a time series filtering task. The state space models we investigated include a standard linear model (Kalman filter) as well as novel non-linear, non-parametric Gaussian Processes (GP) based SSM. We use the AVEC 2014 database for evaluation, which provides ground truth A-V-D labels which allows state and measurement functions to be learned separately simplifying the model training. For the filtering with GP SSM, we used two approximation methods: a recently proposed analytic method and Particle filter. All models were evaluated in terms of average Pearson correlation R and root mean square error (RMSE). The results show that using the same feature vectors, the GP SSMs achieve twice higher correlation and twice smaller RMSE than a Kalman filter.