Emotion spotting: discovering regions of evidence in audio-visual emotion expressions

Emotion spotting: discovering regions of evidence in audio-visual emotion expressions
复制标题

DOI:
10.1145/2993148.2993151
复制
发表时间:
2016-10
期刊:
Proceedings of the 18th ACM International Conference on Multimodal Interaction
影响因子:
--
通讯作者:
Yelin Kim;E. Provost
Yelin Kim;E. Provost
中科院分区:
其他
文献类型:
--
作者:
Yelin Kim;E. Provost

文献摘要

被引文献

相似文献

研究表明,随着时间的推移,人类需要不同数量的信息来准确感知情感表达。这随着情绪等级的不同而有所不同。例如,认识到幸福需要比认识到愤怒需要更长的刺激。然而,以前的自动情感识别系统往往忽略了这些差异。在这项工作中,我们提出了一个数据驱动的框架来探索特定于个体情感类别的情感证据的模式(时间和持续时间)。此外,我们证明了这些模式作为检查通道(下脸、上脸或语音)的函数而不同,并且在不同的实验折叠中出现了一致的模式。我们在情感语料库(IEMOCAP和MSP-IMEV)中也显示了类似的模式。此外,我们还表明,我们提出的方法只使用部分数据(IEMOCAP的准确率为59%),达到了与使用每个话语中的所有数据的系统相当的精度。与随机选择一部分数据的基线方法相比,我们的方法具有更高的精度。我们表明,该方法的性能收益主要来自原型情感表达(定义为具有评分者共识的表达)。这项研究的创新之处在于,它理解了多通道线索是如何随着时间的推移揭示情绪的。
Research has demonstrated that humans require different amounts of information, over time, to accurately perceive emotion expressions. This varies as a function of emotion classes. For example, recognition of happiness requires a longer stimulus than recognition of anger. However, previous automatic emotion recognition systems have often overlooked these differences. In this work, we propose a data-driven framework to explore patterns (timings and durations) of emotion evidence, specific to individual emotion classes. Further, we demonstrate that these patterns vary as a function of which modality (lower face, upper face, or speech) is examined, and consistent patterns emerge across different folds of experiments. We also show similar patterns across emotional corpora (IEMOCAP and MSP-IMPROV). In addition, we show that our proposed method, which uses only a portion of the data (59% for the IEMOCAP), achieves comparable accuracy to a system that uses all of the data within each utterance. Our method has a higher accuracy when compared to a baseline method that randomly chooses a portion of the data. We show that the performance gain of the method is mostly from prototypical emotion expressions (defined as expressions with rater consensus). The innovation in this study comes from its understanding of how multimodal cues reveal emotion over time.