Pre-Finetuning for Few-Shot Emotional Speech Recognition

Pre-Finetuning for Few-Shot Emotional Speech Recognition
复制标题

DOI:
10.48550/arxiv.2302.12921
复制
发表时间:
2023-02
期刊:
--
影响因子:
--
通讯作者:
Maximillian Chen;Zhou Yu
Maximillian Chen;Zhou Yu
中科院分区:
其他
文献类型:
--
作者:
Maximillian Chen;Zhou Yu

文献摘要

相似文献

人们早就知道,语音模型在许多分类任务中会过度拟合单个说话者。这导致扬声器在域外或分布外的设置中泛化不良,这在生产环境中很常见。我们将说话者适应视为一种短时学习问题,并提出研究迁移学习方法的灵感来自于最近在自然语言任务中使用预训练模型的成功。我们提出对困难任务的语音模型进行预微调,以将知识提取到少数几个下游分类目标中。我们对四个多类情绪语音识别语料库的每一个排列进行了预微调,并通过在情绪语音数据集上的33,600次少量微调试验来评估我们的预微调模型。
Speech models have long been known to overfit individual speakers for many classification tasks. This leads to poor generalization in settings where the speakers are out-of-domain or out-of-distribution, as is common in production environments. We view speaker adaptation as a few-shot learning problem and propose investigating transfer learning approaches inspired by recent success with pre-trained models in natural language tasks. We propose pre-finetuning speech models on difficult tasks to distill knowledge into few-shot downstream classification objectives. We pre-finetune Wav2Vec2.0 on every permutation of four multiclass emotional speech recognition corpora and evaluate our pre-finetuned models through 33,600 few-shot fine-tuning trials on the Emotional Speech Dataset.