Sequential Modeling by Leveraging Non-Uniform Distribution of Speech Emotion

Sequential Modeling by Leveraging Non-Uniform Distribution of Speech Emotion
复制标题

DOI:
10.1109/taslp.2023.3244527
复制
发表时间:
2023
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Wei-Cheng Lin;C. Busso
Wei-Cheng Lin;C. Busso
中科院分区:
其他
文献类型:
--
作者:
Wei-Cheng Lin;C. Busso

文献摘要

相似文献

人类情感的表达和感知随着时间的推移并不均匀分布。因此,跟踪片段内情感的局部变化可以产生更好的语音情感识别(SER)模型,即使任务是提供情感内容的句子级预测。探索句子中局部情感变化的一个挑战是,大多数现有情感语料库仅提供句子级注释(即每个句子一个标签)。这种标记方法不适合利用句子中的动态情感趋势。我们提出了一个框架,将句子分成固定数量的块,生成块级情感模式。该方法依靠情感排序来揭示句子中的情感模式,从而创建连续的情感曲线。我们的方法利用检索到的情感曲线,通过序列到序列的公式来训练句子级 SER 模型。所提出的方法在 MSP-Podcast 语料库上实现了唤醒度 (0.7120)、效价 (0.3125) 和支配度 (0.6324) 的最佳一致性相关系数 (CCC) 预测性能。此外,我们通过 IEMOCAP 和 MSP-IMPROV 数据库上的实验验证了该方法。我们进一步将检索到的曲线与时间连续的情绪轨迹进行比较。评估表明,这些检索到的块标签曲线可以有效地捕获句子中的情绪趋势,显示出类似于人类听众注释的时间连续轨迹的时间一致性属性。所提出的 SER 模型学习有意义的、互补的局部信息,有助于改进情感属性的句子级预测。
The expression and perception of human emotions are not uniformly distributed over time. Therefore, tracking local changes of emotion within a segment can lead to better models for speech emotion recognition (SER), even when the task is to provide a sentence-level prediction of the emotional content. A challenge to exploring local emotional changes within a sentence is that most existing emotional corpora only provide sentence-level annotations (i.e., one label per sentence). This labeling approach is not appropriate for leveraging the dynamic emotional trends within a sentence. We propose a framework that splits a sentence into a fixed number of chunks, generating chunk-level emotional patterns. The approach relies on emotion rankers to unveil the emotional pattern within a sentence, creating continuous emotional curves. Our approach trains the sentence-level SER model with a sequence-to-sequence formulation by leveraging the retrieved emotional curves. The proposed method achieves the best concordance correlation coefficient (CCC) prediction performance for arousal (0.7120), valence (0.3125), and dominance (0.6324) on the MSP-Podcast corpus. In addition, we validate the approach with experiments on the IEMOCAP and MSP-IMPROV databases. We further compare the retrieved curves with time-continuous emotional traces. The evaluation demonstrates that these retrieved chunk-label curves can effectively capture emotional trends within a sentence, displaying a time-consistency property that is similar to time-continuous traces annotated by human listeners. The proposed SER model learns meaningful, complementary, local information that contributes to the improvement of sentence-level predictions of emotional attributes.