Multimodal continuous affect recognition based on LSTM and multiple kernel learning

Multimodal continuous affect recognition based on LSTM and multiple kernel learning
复制标题

DOI:
10.1109/apsipa.2014.7041743
复制
发表时间:
2014-12
期刊:
Signal and Information Processing Association Annual Summit and Conference (APSIPA), 2014 Asia-Pacific
影响因子:
--
通讯作者:
Jiamei Wei;Ercheng Pei;D. Jiang;H. Sahli;Lei Xie;Zhonghua Fu
Jiamei Wei;Ercheng Pei;D. Jiang;H. Sahli;Lei Xie;Zhonghua Fu
中科院分区:
其他
文献类型:
--
作者:
Jiamei Wei;Ercheng Pei;D. Jiang;H. Sahli;Lei Xie;Zhonghua Fu

文献摘要

被引文献

相似文献

本文提出了一种基于长短期记忆递归神经网络(LSTM-RNN)和多核学习(MKL)的多模态情感识别方案(LSTM-MKL)。它利用LSTM-RNN的优势来建模连续观测之间的长期依赖关系,并使用MKL功率来建模输入和输出之间的非线性相关性。对于每个影响维度(唤醒、效价、期望和力量),训练了两个LSTM-RNN模型,每个模型一个。在识别阶段,音频和视觉特征被输入到相应的学习LSTM模型中,该模型反过来产生对影响维度的初始估计。LSTM输出进一步输入到多核支持向量回归(MK-SVR)中进行最终识别。在AVEC2012数据库上进行的实验结果表明,与传统的SVR-LLR(支持向量机-局部线性回归)或MK-SVR融合方案相比,LSTM-MKL融合方案获得了更高的识别效果,其相关系数(COR)为0.354,而SVR-LLR和MK-SVR的相关系数分别为0.124和0.168。
In this paper, we propose a Long Short-Term Memory Recurrent Neural Network (LSTM-RNN) and multiple kernel learning (MKL) based multi-modal affect recognition scheme (LSTM-MKL). It takes the LSTM-RNN advantage to model the long range dependencies between successive observations, and uses the MKL power to model the non-linear correlations between the inputs and outputs. For each of the affect dimensions (arousal, valence, expectancy, and power), two LSTM-RNN models are trained, one for each modality. In the recognition phase, the audio and visual features are input to the corresponding learned LSTM models, which in turn produce initial estimates of the affect dimensions. The LSTM outputs are further input into a multi-kernel support vector regression (MK-SVR) for the final recognition. Experimental results carried out on the AVEC2012 database, show that compared to the traditional SVR-LLR (Support Vector Machine - local linear regression) or MK-SVR fusion scheme, the proposed LSTM-MKL fusion scheme obtains higher recognition results, with an correlation coefficient (COR) of 0.354, compared to a COR of 0.124 for SVR-LLR, and 0.168 for MK-SVR, respectively.