Continual Learning for Monolingual End-to-End Automatic Speech Recognition

Continual Learning for Monolingual End-to-End Automatic Speech Recognition
复制标题

单语端到端自动语音识别的持续学习

DOI:
--
复制
发表时间:
2021
期刊:
European Signal Processing Conference
影响因子:
--
通讯作者:
H. V. hamme
H. V. hamme
中科院分区:
--
文献类型:
--
作者:
Steven Vander Eeckt;H. V. hamme

文献摘要

被引文献

相似文献

自动语音识别(ASR)模型适应新的领域会导致原始领域的性能下降,这种现象称为灾难性遗忘(CF)。即使是单语ASR模型也无法扩展到新的口音、方言、主题等,而不会受到CF的影响,这使得它们在不存储所有过去数据的情况下无法持续增强。幸运的是,可以使用持续学习(CL)方法,其目的是在克服CF的同时实现持续适应。在本文中,我们为端到端ASR实现了大量的CL方法,并测试和比较了它们在四个新任务中扩展单语混合CTC-Transformer模型的能力。我们发现,性能最好的CL方法将微调模型(下限)和所有任务联合训练的模型(上限)之间的差距缩小了40%以上,同时只需要访问0.6%的原始数据。
Adapting Automatic Speech Recognition (ASR) models to new domains results in a deterioration of performance on the original domain(s), a phenomenon called Catastrophic Forgetting (CF). Even monolingual ASR models cannot be extended to new accents, dialects, topics, etc. without suffering from CF, making them unable to be continually enhanced without storing all past data. Fortunately, Continual Learning (CL) methods, which aim to enable continual adaptation while overcoming CF, can be used. In this paper, we implement an extensive number of CL methods for End-to-End ASR and test and compare their ability to extend a monolingual Hybrid CTC-Transformer model across four new tasks. We find that the best performing CL method closes the gap between the fine-tuned model (lower bound) and the model trained jointly on all tasks (upper bound) by more than 40%, while requiring access to only 0.6% of the original data.