Predicting the distribution of emotion perception: capturing inter-rater variability

Predicting the distribution of emotion perception: capturing inter-rater variability
复制标题

DOI:
10.1145/3136755.3136792
复制
发表时间:
2017-11
期刊:
Proceedings of the 19th ACM International Conference on Multimodal Interaction
影响因子:
--
通讯作者:
Biqiao Zhang;Georg Essl;E. Provost
Biqiao Zhang;Georg Essl;E. Provost
中科院分区:
其他
文献类型:
--
作者:
Biqiao Zhang;Georg Essl;E. Provost

文献摘要

相似文献

情绪感知是因人而异的。情感的维度表征可以通过根据其属性(例如,效价,正对负,以及激活,平静对兴奋)。然而,在许多情感识别系统中,这种可变性通常被认为是“噪声”,并通过跨评分者进行平均来衰减。然而,评分者间的差异性提供了有关情绪表达的微妙或清晰的信息,并可用于描述复杂的情绪。在本文中,我们研究的方法,可以有效地捕捉变异的评价者通过预测情绪感知的离散概率分布的效价激活空间。我们建议:(1)一种标签处理方法,可以从有限数量的有序标签中生成情感的二维离散概率分布;(2)一种新的方法,使用动态视听特征和卷积神经网络(CNN)预测生成的概率分布。在MSP-IMPROV语料库上的实验结果表明,该方法比传统的支持向量回归(SVR)方法更有效,具有话语级统计特征,并且音频和视频模态的特征级融合优于决策级融合。所提出的CNN模型主要提高了价维度的预测精度,并带来了与自然相互作用记录的数据一致的性能改善。结果表明,从有限数量的标签生成情感分布,并使用动态特征和神经网络预测分布的有效性。
Emotion perception is person-dependent and variable. Dimensional characterizations of emotion can capture this variability by describing emotion in terms of its properties (e.g., valence, positive vs. negative, and activation, calm vs. excited). However, in many emotion recognition systems, this variability is often considered "noise" and is attenuated by averaging across raters. Yet, inter-rater variability provides information about the subtlety or clarity of an emotional expression and can be used to describe complex emotions. In this paper, we investigate methods that can effectively capture the variability across evaluators by predicting emotion perception as a discrete probability distribution in the valence-activation space. We propose: (1) a label processing method that can generate two-dimensional discrete probability distributions of emotion from a limited number of ordinal labels; (2) a new approach that predicts the generated probabilistic distributions using dynamic audio-visual features and Convolutional Neural Networks (CNNs). Our experimental results on the MSP-IMPROV corpus suggest that the proposed approach is more effective than the conventional Support Vector Regressions (SVRs) approach with utterance-level statistical features, and that feature-level fusion of the audio and video modalities outperforms decision-level fusion. The proposed CNN model predominantly improves the prediction accuracy for the valence dimension and brings a consistent performance improvement over data recorded from natural interactions. The results demonstrate the effectiveness of generating emotion distributions from limited number of labels and predicting the distribution using dynamic features and neural networks.