Domain Expansion for End-to-End Speech Recognition: Applications for Accent/Dialect Speech

Domain Expansion for End-to-End Speech Recognition: Applications for Accent/Dialect Speech
复制标题

DOI:
10.1109/taslp.2022.3233238
复制
发表时间:
2023
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Shahram Ghorbani;J. Hansen
Shahram Ghorbani;J. Hansen
中科院分区:
其他
文献类型:
--
作者:
Shahram Ghorbani;J. Hansen

文献摘要

相似文献

训练自动语音识别(ASR)系统与顺序输入的数据从交替域是一个重要的里程碑,以达到人类的可懂度水平的语音识别。顺序学习的主要挑战是,目前的适应技术导致以前看到的领域的显着性能下降。为了缓解灾难性遗忘问题,本研究提出了两种情况下有效的域扩展技术:1)只有新的域数据可用,2)现有和新的域数据都可用。我们通过实验来检验这些方法的有效性,实验是用母语英语训练的模型来适应不同的英语口音。对于第一种情况,我们研究了几种现有的和提出的基于正则化的方法,以减轻初始数据的性能损失。实验表明,我们提出的软KL发散(SKLD)模型平均(MA)方法的上级性能。在这种方法中,SKLD首先解决了自适应过程中的遗忘问题;接下来,MA通过对初始模型和自适应模型的参数进行平均,在两个域之间进行最终的有效折衷。对于第二种情况,我们探索了几种基于排练的方法,这些方法利用初始数据来保持原始模型的性能。我们提出了梯度平均(GA)以及一种方法,该方法通过对初始域和新域计算的梯度进行平均来操作。实验表明,GA优于再训练和专门设计的持续学习方法,如平均梯度情节记忆(AGEM)。此外,GA显着提高了计算成本的完整的再训练方法。
Training Automatic Speech Recognition (ASR) systems with sequentially incoming data from alternate domains is an essential milestone in order to reach human intelligibility level in speech recognition. The main challenge of sequential learning is that current adaptation techniques result in significant performance degradation for previously-seen domains. To mitigate the catastrophic forgetting problem, this study proposes effective domain expansion techniques for two scenarios: 1) where only new domain data is available, and 2) where both prior and new domain data are available. We examine the efficacy of the approaches through experiments on adapting a model trained with native English to different English accents. For the first scenario, we study several existing and proposed regularization-based approaches to mitigate performance loss of initial data. The experiments demonstrate the superior performance of our proposed Soft KL-Divergence (SKLD)-Model Averaging (MA) approach. In this approach, SKLD first alleviates the forgetting problem during adaptation; next, MA makes the final efficient compromise between the two domains by averaging parameters of the initial and adapted models. For the second scenario, we explore several rehearsal-based approaches, which leverage initial data to maintain the original model performance. We propose Gradient Averaging (GA) as well as an approach which operates by averaging gradients computed for both initial and new domains. Experiments demonstrate that GA outperforms retraining and specifically designed continual learning approaches, such as Averaged Gradient Episodic Memory (AGEM). Moreover, GA significantly improves computational costs over the complete retraining approach.