Nonnegative Matrix Factorization Based Transfer Subspace Learning for Cross-Corpus Speech Emotion Recognition

Nonnegative Matrix Factorization Based Transfer Subspace Learning for Cross-Corpus Speech Emotion Recognition
复制标题

基于非负矩阵分解的跨语料库语音情感识别的迁移子空间学习

DOI:
10.1109/taslp.2020.3006331
复制
发表时间:
2020-01-01
影响因子:
5.4
通讯作者:
Han, Jiqing
Han, Jiqing
中科院分区:
计算机科学2区
文献类型:
--
作者:
Luo, Hui;Han, Jiqing

文献摘要

被引文献

相似文献

本文主要研究跨语料库语音情感识别问题。针对训练样本和测试样本分布不一致的问题,提出了一种基于非负矩阵分解的迁移子空间学习方法(NMFTSL).该方法试图为源语料库和目标语料库找到一个共享的特征子空间,在该子空间中尽可能消除源语料库和目标语料库之间的差异,并排除源语料库和目标语料库中的个别成分,从而将源语料库中的知识转移到目标语料库中。具体来说,在这个诱导子空间中,我们不仅最小化边缘分布之间的距离,而且最小化条件分布之间的距离,其中这两个距离都是由最大平均差异准则测量的。为了估计目标语料的条件分布,我们提出将目标标签的预测和特征表示的学习集成到一个联合学习模型中。同时,我们引入了一个差异损失来排除共享子空间中的个体分量,这可以进一步减少源和目标个体分量之间的相互干扰。此外,我们提出了一个歧视损失的标签引入到共享子空间,这可以提高的歧视能力的特征表示。并给出了相应优化问题的求解方法。为了评估我们的方法的性能,我们使用6个流行的语音情感语料库构建了30个跨语料库SER方案。实验结果表明,我们的方法实现了更好的整体性能比国家的最先进的方法。
This article focuses on the cross-corpus speech emotion recognition (SER) task. To overcome the problem that the distribution of training (source) samples is inconsistent with that of testing (target) samples, we propose a non-negative matrix factorization based transfer subspace learning method (NMFTSL). Our method tries to find a shared feature subspace for the source and target corpora, in which the discrepancy between the two distributions is eliminated as much as possible and their individual components are excluded, thus the knowledge of the source corpus can be transferred to the target corpus. Specifically, in this induced subspace, we minimize the distances not only between the marginal distributions but also between the conditional distributions, where both distances are measured by the maximum mean discrepancy criterion. To estimate the conditional distribution of the target corpus, we propose to integrate the prediction of target label and the learning of feature representation into a joint learning model. Meanwhile, we introduce a difference loss to exclude the individual components from the shared subspace, which can further reduce the mutual interference between the source and target individual components. Moreover, we propose a discrimination loss to introduce the labels into the shared subspace, which can improve the discrimination ability of the feature representation. We also provide the solution for the corresponding optimization problem. To evaluate the performance of our method, we construct 30 cross-corpus SER schemes using 6 popular speech emotion corpora. Experimental results show that our approach achieves better overall performance than state-of-the-art methods.