Best Practices for Noise-Based Augmentation to Improve the Performance of Emotion Recognition "In the Wild"

Best Practices for Noise-Based Augmentation to Improve the Performance of Emotion Recognition "In the Wild"
复制标题

基于噪声的增强的最佳实践,以提高“野外”情绪识别的性能

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
E. Provost
E. Provost
中科院分区:
--
文献类型:
--
作者:
Mimansa Jaiswal;E. Provost

文献摘要

参考文献

被引文献

相似文献

情感识别作为高风险下游应用的关键组成部分已被证明是有效的,例如课堂参与或心理健康评估。这些系统通常在单个实验室环境中收集的小数据集上进行训练,因此在测试具有不同噪声特征的数据时会出现问题。已经提出了多种基于噪声的数据增强方法来应对其他语音域中的这一挑战。但是,与语音识别和说话人验证不同,在情感识别中,基于噪声的数据增强可能会改变原始情感样本的底层标签。在这项工作中,我们产生现实的噪声样本的一个众所周知的情感数据集(IEMOCAP)使用多种类别的环境和合成噪声。我们评估了当噪声引入时,人类和机器的情感感知如何变化。我们发现,一些常用的情感识别增强技术会显著改变人类的感知,这可能导致不可靠的评估指标,例如评估对抗性攻击的效率。我们还发现,经过训练的最先进的情感识别模型无法对看不见的噪声增强样本进行分类,即使在噪声增强数据集上进行训练。这一发现表明了这些系统在现实世界条件下的脆弱性。我们提出了一组基于噪声的情感数据集增强的建议,以及如何部署这些情感识别系统“在野外”。
Emotion recognition as a key component of high-stake downstream applications has been shown to be effective, such as classroom engagement or mental health assessments. These systems are generally trained on small datasets collected in single laboratory environments, and hence falter when tested on data that has different noise characteristics. Multiple noise-based data augmentation approaches have been proposed to counteract this challenge in other speech domains. But, unlike speech recognition and speaker verification, in emotion recognition, noise-based data augmentation may change the underlying label of the original emotional sample. In this work, we generate realistic noisy samples of a well known emotion dataset (IEMOCAP) us-ing multiple categories of environmental and synthetic noise. We evaluate how both human and machine emotion perception changes when noise is introduced. We find that some commonly used augmentation techniques for emotion recognition significantly change human perception, which may lead to unreliable evaluation metrics such as evaluating ef-ficiency of adversarial attack. We also find that the trained state-of-the-art emotion recognition models fail to classify unseen noise-augmented samples, even when trained on noise augmented datasets. This finding demonstrates the brittleness of these systems in real-world conditions. We propose a set of recommendations for noise-based augmentation of emotion datasets and for how to deploy these emotion recognition systems “in the wild”.
DOI: 10.1109/taslp.2018.2867099
发表时间: 2018-12-01
影响因子: 5.4
作者:
Abdelwahab, Mohammed;Busso, Carlos
通讯作者: Busso, Carlos
AllenNLP Interpret:解释 NLP 模型预测的框架
DOI: 10.18653/v1/d19-3002
发表时间: 2019
期刊: Conference on Empirical Methods in Natural Language Processing (EMNLP
影响因子: --
作者:
Wallace, Eric;Tuyls, Jens;Wang, Junlin;Subramanian, Sanjay;Gardner, Matt;Singh, Sameer
通讯作者: Singh, Sameer
DOI: 10.1016/0005-7916(94)90063-9
发表时间: 1994-03-01
影响因子: 1.8
作者:
BRADLEY, MM;LANG, PJ
通讯作者: LANG, PJ