Representation Learning Through Cross-Modal Conditional Teacher-Student Training For Speech Emotion Recognition

Representation Learning Through Cross-Modal Conditional Teacher-Student Training For Speech Emotion Recognition
复制标题

通过跨模态条件师生训练进行表征学习以实现语音情感识别

DOI:
10.1109/icassp43922.2022.9747754
复制
发表时间:
2021
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
K. Kirchhoff
K. Kirchhoff
中科院分区:
--
文献类型:
--
作者:
S. Srinivasan;Zhaocheng Huang;K. Kirchhoff

文献摘要

参考文献

被引文献

相似文献

通用的预训练语音和文本表示有望减少特定语音和语言任务对大型标记数据集的需求。然而,目前尚不清楚如何有效地将这些表示应用于语音情感识别。最近的公共基准显示了几种流行的自我监督语音表示在情感分类方面的功效。在这项研究中,我们表明,表现最好的表示之间的主要差异在于预测效价,而预测激活和优势维度方面的差异不太明显。然而,我们表明,与也包含文本表示的多模态模型相比,即使是性能最好的 HuBERT 表示在价预测方面也表现不佳。我们通过使用多模态模型作为教师将词汇信息注入语音表示来解决这个缺点。为了提高我们方法的有效性,我们提出了一种对情绪预测质量的新颖估计,以适应师生培训。我们报告了 MSP-Podcast 语料库上新的纯音频最先进的激活、效价和优势预测的一致性相关系数 (CCC) 值分别为 0.757、0.627、0.671,并且在 IEMOCAP 语料库上的最先进值分别为 0.667、0.582、0.545。
Generic pre-trained speech and text representations promise to reduce the need for large labeled datasets on specific speech and language tasks. However, it is not clear how to effectively adapt these representations for speech emotion recognition. Recent public benchmarks show the efficacy of several popular self-supervised speech representations for emotion classification. In this study, we show that the primary difference between the top-performing representations is in predicting valence while the differences in predicting activation and dominance dimensions are less pronounced. However, we show that even the best-performing HuBERT representation underperforms on valence prediction compared to a multimodal model that also incorporates text representation. We address this shortcoming by injecting lexical information into the speech representation using the multimodal model as a teacher. To improve the efficacy of our approach, we propose a novel estimate of the quality of the emotion predictions, to condition teacher-student training. We report new audio-only state-of-the-art concordance correlation coefficient (CCC) values of 0.757, 0.627, 0.671 for activation, valence and dominance predictions, respectively, on the MSP-Podcast corpus, and also state-of-the-art values of 0.667, 0.582, 0.545 on the IEMOCAP corpus.
DOI: 10.1109/asru51503.2021.9688093
发表时间: 2021-07
期刊: 2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
影响因子: --
作者:
Ankita Pasad;Ju-Chieh Chou;Karen Livescu
通讯作者: Ankita Pasad;Ju-Chieh Chou;Karen Livescu