Error Correction in ASR using Sequence-to-Sequence Models

Error Correction in ASR using Sequence-to-Sequence Models
复制标题

使用序列到序列模型进行 ASR 纠错

DOI:
--
复制
发表时间:
2022
期刊:
arXiv.org
影响因子:
--
通讯作者:
P. Jyothi
P. Jyothi
中科院分区:
--
文献类型:
--
作者:
S. Dutta;Shreyansh Jain;Ayush Maheshwari;Ganesh Ramakrishnan;P. Jyothi

文献摘要

被引文献

相似文献

自动语音识别(ASR)中的后期编辑需要自动纠正由ASR系统产生的常见和系统错误。ASR系统的输出在很大程度上容易出现语音和拼写错误。在本文中,我们建议使用一个强大的预先训练的序列到序列模型BART,并进一步自适应地训练作为去噪模型,来纠正这种类型的错误。自适应训练在通过综合诱导误差以及通过结合来自现有ASR系统的实际误差而获得的扩充数据集上执行。我们还提出了一种简单的方法来使用词级对齐来对输出进行重新评分。在有重音语音数据上的实验结果表明,我们的策略有效地纠正了大量的ASR错误,并且与竞争基准相比,得到了更好的WER结果。我们还强调了在印地语相关语法纠错任务中获得的否定结果,表明了我们所提出的模型在捕捉更广泛的上下文方面的局限性。
Post-editing in Automatic Speech Recognition (ASR) entails automatically correcting common and systematic errors produced by the ASR system. The outputs of an ASR system are largely prone to phonetic and spelling errors. In this paper, we propose to use a powerful pre-trained sequence-to-sequence model, BART, further adaptively trained to serve as a denoising model, to correct errors of such types. The adaptive training is performed on an augmented dataset obtained by synthetically inducing errors as well as by incorporating actual errors from an existing ASR system. We also propose a simple approach to rescore the outputs using word level alignments. Experimental results on accented speech data demonstrate that our strategy effectively rectifies a significant number of ASR errors and produces improved WER results when compared against a competitive baseline. We also highlight a negative result obtained on the related grammatical error correction task in Hindi language showing the limitation in capturing wider context by our proposed model.