On-the-Fly Aligned Data Augmentation for Sequence-to-Sequence ASR

On-the-Fly Aligned Data Augmentation for Sequence-to-Sequence ASR
复制标题

DOI:
10.21437/interspeech.2021-1679
复制
发表时间:
2021-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Tsz Kin Lam;Mayumi Ohta;Shigehiko Schamoni;S. Riezler
Tsz Kin Lam;Mayumi Ohta;Shigehiko Schamoni;S. Riezler
中科院分区:
其他
文献类型:
--
作者:
Tsz Kin Lam;Mayumi Ohta;Shigehiko Schamoni;S. Riezler

文献摘要

被引文献

相似文献

我们提出了一种用于自动语音识别(ASR)的动态数据增强方法,该方法使用对齐信息来生成有效的训练样本。我们的方法称为ASR的对齐数据增强(ADA),以对齐的方式替换转录的标记和语音表示,以生成以前看不见的训练对。语音表示从已经从训练语料库中提取的音频词典中采样,并将说话者变化注入训练示例中。转录的标记由语言模型预测,使得增强的数据对在语义上接近原始数据,或者随机采样。这两种策略都会产生训练对,提高ASR训练的鲁棒性。我们在Seq-to-Seq架构上的实验表明,ADA可以应用于SpecAugment之上,并在LibriSpeech 100 h和LibriSpeech 960 h测试数据集上分别实现了比SpecAugment单独WER约9-23%和4-15%的相对改进。
We propose an on-the-fly data augmentation method for automatic speech recognition (ASR) that uses alignment information to generate effective training samples. Our method, called Aligned Data Augmentation (ADA) for ASR, replaces transcribed tokens and the speech representations in an aligned manner to generate previously unseen training pairs. The speech representations are sampled from an audio dictionary that has been extracted from the training corpus and inject speaker variations into the training examples. The transcribed tokens are either predicted by a language model such that the augmented data pairs are semantically close to the original data, or randomly sampled. Both strategies result in training pairs that improve robustness in ASR training. Our experiments on a Seq-to-Seq architecture show that ADA can be applied on top of SpecAugment, and achieves about 9-23% and 4-15% relative improvements in WER over SpecAugment alone on LibriSpeech 100h and LibriSpeech 960h test datasets, respectively.