Towards Noise Robust Speech Emotion Recognition Using Dynamic Layer Customization
Towards Noise Robust Speech Emotion Recognition Using Dynamic Layer Customization
复制标题
DOI:
10.1109/acii52823.2021.9597437
复制
发表时间:
2021-09
期刊:
影响因子:
--
通讯作者:
Alex Wilf;E. Provost
中科院分区:
文献类型:
--
作者:
Alex Wilf;E. Provost
Robustness to environmental noise is important to creating automatic speech emotion recognition systems that are deployable in the real world. In this work, we experiment with two paradigms, one where we can anticipate noise sources that will be seen at test time and one where we cannot. In our first experiment, we assume that we have advance knowledge of the noise conditions that will be seen at test time. We show that we can use this knowledge to create "expert" feature encoders for each noise condition. If the noise condition is unchanging, data can be routed to a single encoder to improve robustness. However, if the noise source is variant, this paradigm is too restrictive. In-stead, we introduce a new approach, dynamic layer customization (DLC), that allows the data to be dynamically routed to noise-matched encoders and then recombined. Critically, this process maintains temporal order, enabling extensions for multimodal models that generally benefit from long-term context. In our second experiment, we investigate whether partial knowledge of noise seen at test time can still be used to train systems that generalize well to unseen noise conditions using state-of-the-art domain adaptation algorithms. We find that DLC enables performance increases in both cases, highlighting the utility of mixture-of-expert approaches, domain adaptation methods and DLC to noise robust automatic speech emotion recognition.