Monologue versus Conversation: Differences in Emotion Perception and Acoustic Expressivity

Monologue versus Conversation: Differences in Emotion Perception and Acoustic Expressivity
复制标题

DOI:
10.1109/acii55700.2022.9953814
复制
发表时间:
2022-10
期刊:
2022 10th International Conference on Affective Computing and Intelligent Interaction (ACII)
影响因子:
--
通讯作者:
Woan-Shiuan Chien;Shreya G. Upadhyay;Wei-Cheng Lin;Ya-Tse Wu;Bo-Hao Su;C. Busso;Chi-Chun Lee
Woan-Shiuan Chien;Shreya G. Upadhyay;Wei-Cheng Lin;Ya-Tse Wu;Bo-Hao Su;C. Busso;Chi-Chun Lee
中科院分区:
其他
文献类型:
--
作者:
Woan-Shiuan Chien;Shreya G. Upadhyay;Wei-Cheng Lin;Ya-Tse Wu;Bo-Hao Su;C. Busso;Chi-Chun Lee

文献摘要

相似文献

推进语音情感识别(SER)高度依赖于用于训练模型的源,即情感语音语料库。通过排列不同的设计参数,研究人员已经发布了一些版本的语料库,这些语料库试图为训练SER提供更高质量的资源。在这项工作中,我们重点研究了收藏的传播模式。特别是,我们分析在人际对话或独白中收集的情感语言模式。虽然大家都知道,对话为激发真实的情感表达提供了更好的协议,但缺乏系统的分析来确定对话语音是否提供了“更好质量”的来源。具体来说,我们从三个角度来研究这个研究问题:感知差异、声学变异性和SER模型学习。我们对MSP-Podcast语料库的分析表明:1)在评估分类情绪时,评分者对对话录音的一致性更高;2)在对话中观察到的感知和声学模式具有与情感文献中讨论的预期趋势更一致的属性;3)可以从会话数据中训练出更健壮的SER模型。这项工作带来了初步的证据,表明对话样本可能比独白样本提供更好的质量来源,用于构建SER模型。
Advancing speech emotion recognition (SER) depends highly on the source used to train the model, i.e., the emotional speech corpora. By permuting different design parameters, researchers have released versions of corpora that attempt to provide a better-quality source for training SER. In this work, we focus on studying communication modes of collection. In particular, we analyze the patterns of emotional speech collected during interpersonal conversations or monologues. While it is well known that conversation provides a better protocol for eliciting authentic emotion expressions, there is a lack of systematic analyses to determine whether conversational speech provide a “better-quality” source. Specifically, we examine this research question from three perspectives: perceptual differences, acoustic variability and SER model learning. Our analyses on the MSP-Podcast corpus show that: 1) rater's consistency for conversation recordings is higher when evaluating categorical emotions, 2) the perceptions and acoustic patterns observed on conversations have properties that are better aligned with expected trends discussed in emotion literature, and 3) a more robust SER model can be trained from conversational data. This work brings initial evidences stating that samples of conversations may provide a better-quality source than samples from monologues for building a SER model.