Towards Noise Robust Speech Emotion Recognition Using Dynamic Layer Customization

Towards Noise Robust Speech Emotion Recognition Using Dynamic Layer Customization
复制标题

DOI:
10.1109/acii52823.2021.9597437
复制
发表时间:
2021-09
期刊:
2021 9th International Conference on Affective Computing and Intelligent Interaction (ACII)
影响因子:
--
通讯作者:
Alex Wilf;E. Provost
Alex Wilf;E. Provost
中科院分区:
其他
文献类型:
--
作者:
Alex Wilf;E. Provost

文献摘要

相似文献

对环境噪声的鲁棒性对于创建可在现实世界中部署的自动语音情感识别系统非常重要。在这项工作中,我们尝试了两种范例,一种是我们可以预测测试时会看到的噪声源,另一种是我们无法预测的。在我们的第一个实验中,我们假设我们预先了解测试时出现的噪声条件。我们表明,我们可以利用这些知识为每种噪声条件创建“专家”特征编码器。如果噪声条件不变,则可以将数据路由到单个编码器以提高鲁棒性。然而,如果噪声源是变化的,则该范例的限制性太大。相反,我们引入了一种新方法,即动态层定制(DLC),它允许将数据动态路由到噪声匹配编码器,然后重新组合。至关重要的是,这个过程保持了时间顺序,从而能够扩展通常受益于长期背景的多模态模型。在我们的第二个实验中,我们研究了在测试时看到的噪声的部分知识是否仍然可以用于训练使用最先进的域适应算法很好地泛化到未见噪声条件的系统。我们发现 DLC 在这两种情况下都能提高性能,突出了专家混合方法、域适应方法和 DLC 对噪声稳健的自动语音情感识别的实用性。
Robustness to environmental noise is important to creating automatic speech emotion recognition systems that are deployable in the real world. In this work, we experiment with two paradigms, one where we can anticipate noise sources that will be seen at test time and one where we cannot. In our first experiment, we assume that we have advance knowledge of the noise conditions that will be seen at test time. We show that we can use this knowledge to create "expert" feature encoders for each noise condition. If the noise condition is unchanging, data can be routed to a single encoder to improve robustness. However, if the noise source is variant, this paradigm is too restrictive. In-stead, we introduce a new approach, dynamic layer customization (DLC), that allows the data to be dynamically routed to noise-matched encoders and then recombined. Critically, this process maintains temporal order, enabling extensions for multimodal models that generally benefit from long-term context. In our second experiment, we investigate whether partial knowledge of noise seen at test time can still be used to train systems that generalize well to unseen noise conditions using state-of-the-art domain adaptation algorithms. We find that DLC enables performance increases in both cases, highlighting the utility of mixture-of-expert approaches, domain adaptation methods and DLC to noise robust automatic speech emotion recognition.