Exploring Wav2vec 2.0 Fine Tuning for Improved Speech Emotion Recognition

Exploring Wav2vec 2.0 Fine Tuning for Improved Speech Emotion Recognition
复制标题

探索 Wav2vec 2.0 微调以改进语音情感识别

DOI:
--
复制
发表时间:
2021
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
Alexander I. Rudnicky
Alexander I. Rudnicky
中科院分区:
--
文献类型:
--
作者:
Li;Alexander I. Rudnicky

文献摘要

被引文献

相似文献

虽然Wav 2 Vec 2.0已经被提出用于语音识别(ASR),但它也可以用于语音情感识别(SER);使用不同的微调策略可以显着提高其性能。首先提出了两种基本方法,香草微调(V-FT)和任务自适应预训练(TAPT)。我们表明,V-FT能够在IEMOCAP数据集上超越最先进的模型。TAPT,现有的NLP微调策略,进一步提高SER的性能。我们还介绍了一种新的微调方法,称为P-TAPT,它修改了TAPT目标学习情境化的情感表示。实验表明,P-TAPT的性能优于TAPT,特别是在低资源设置。与本文献中的先前作品相比,我们的顶线系统在IEMOCAP上的最先进性能上实现了7.4%的未加权准确度(UA)的绝对改善。我们的代码是公开的。1
While Wav2Vec 2.0 has been proposed for speech recognition (ASR), it can also be used for speech emotion recognition (SER); its performance can be significantly improved using different fine-tuning strategies. Two baseline methods, vanilla fine-tuning (V-FT) and task adaptive pretraining (TAPT) are first presented. We show that V-FT is able to outperform state-of-the-art models on the IEMOCAP dataset. TAPT, an existing NLP fine-tuning strategy, further improves the performance on SER. We also introduce a novel fine-tuning method termed P-TAPT, which modifies the TAPT objective to learn contextualized emotion representations. Experiments show that P-TAPT performs better than TAPT, especially under low-resource settings. Compared to prior works in this literature, our top-line system achieved a 7.4% absolute improvement in unweighted accuracy (UA) over the state-of-the-art performance on IEMOCAP. Our code is publicly available.1