Multiview Supervised Dictionary Learning in Speech Emotion Recognition

Multiview Supervised Dictionary Learning in Speech Emotion Recognition
复制标题

DOI:
10.1109/taslp.2014.2319157
复制
发表时间:
2014-06
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
M. Gangeh;Pouria Fewzee;A. Ghodsi;M. Kamel;F. Karray
M. Gangeh;Pouria Fewzee;A. Ghodsi;M. Kamel;F. Karray
中科院分区:
其他
文献类型:
--
作者:
M. Gangeh;Pouria Fewzee;A. Ghodsi;M. Kamel;F. Karray

文献摘要

被引文献

相似文献

最近,一种基于希尔伯特-施密特独立性准则(HSIC)的监督字典学习(SDL)方法已经被提出,该方法在数据和相应标签之间的依赖性最大化的空间中学习字典和相应的稀疏系数。在本文中,两个多视图字典学习技术提出了基于HSIC的SDL。虽然这两种技术中的一种学习一个字典和所有视图中融合特征的空间中的对应系数,但另一种在每个视图中学习一个字典,随后在学习的字典的空间中融合稀疏系数。在语音情感识别(SER)的应用中,证明了所提出的多视图学习技术在使用单视图的互补信息的有效性。AVEC 2012数据集的全连续子挑战(FCSC)用于两个不同的视图:基线和光谱能量分布(SED)特征集。四维影响,即,唤醒,期望,权力,和效价预测使用所提出的多视图方法作为连续响应变量。将结果与单视图、AVEC 2012基线系统以及文献中的其他监督和无监督多视图学习方法进行了比较。使用相关系数作为性能指标,在预测的连续尺寸的影响,它表明,所提出的方法实现了最高的性能之间的竞争对手。两个建议的多视图技术的相对性能和它们之间的关系进行了讨论。特别是,它表明,通过提供一个额外的约束,这些方法之一的字典,它变得相同的其他。
Recently, a supervised dictionary learning (SDL) approach based on the Hilbert-Schmidt independence criterion (HSIC) has been proposed that learns the dictionary and the corresponding sparse coefficients in a space where the dependency between the data and the corresponding labels is maximized. In this paper, two multiview dictionary learning techniques are proposed based on this HSIC-based SDL. While one of these two techniques learns one dictionary and the corresponding coefficients in the space of fused features in all views, the other learns one dictionary in each view and subsequently fuses the sparse coefficients in the spaces of learned dictionaries. The effectiveness of the proposed multiview learning techniques in using the complementary information of single views is demonstrated in the application of speech emotion recognition (SER). The fully-continuous sub-challenge (FCSC) of the AVEC 2012 dataset is used in two different views: baseline and spectral energy distribution (SED) feature sets. Four dimensional affects, i.e., arousal, expectation, power, and valence are predicted using the proposed multiview methods as the continuous response variables. The results are compared with the single views, AVEC 2012 baseline system, and also other supervised and unsupervised multiview learning approaches in the literature. Using correlation coefficient as the performance measure in predicting the continuous dimensional affects, it is shown that the proposed approach achieves the highest performance among the rivals. The relative performance of the two proposed multiview techniques and their relationship are also discussed. Particularly, it is shown that by providing an additional constraint on the dictionary of one of these approaches, it becomes the same as the other.