N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space

N-best T5: Robust ASR Error Correction using Multiple Input Hypotheses and Constrained Decoding Space
复制标题

DOI:
10.21437/interspeech.2023-1616
复制
发表时间:
2023-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Rao Ma;M. Gales;K. Knill;Mengjie Qian
Rao Ma;M. Gales;K. Knill;Mengjie Qian
中科院分区:
其他
文献类型:
--
作者:
Rao Ma;M. Gales;K. Knill;Mengjie Qian

文献摘要

被引文献

相似文献

纠错模型构成了自动语音识别(ASR)后处理的重要组成部分,以提高转录的可读性和质量。大多数以前的作品使用1-Best ASR假设作为输入,因此只能通过利用一个句子内的上下文来进行更正。在这项工作中,我们提出了一种新的N-Best T5模型,该模型在T5模型的基础上进行了微调,并使用ASR N-Best列表作为模型输入。通过传递来自预训练语言模型的知识和从ASR解码空间获得更丰富的信息,所提出的方法的性能优于强Conformer-Transducer基线。标准纠错的另一个问题是生成过程没有得到很好的指导。为了解决这一问题,使用了基于N最佳列表或ASR点阵的受限解码过程,其允许传播附加信息。
Error correction models form an important part of Automatic Speech Recognition (ASR) post-processing to improve the readability and quality of transcriptions. Most prior works use the 1-best ASR hypothesis as input and therefore can only perform correction by leveraging the context within one sentence. In this work, we propose a novel N-best T5 model for this task, which is fine-tuned from a T5 model and utilizes ASR N-best lists as model input. By transferring knowledge from the pre-trained language model and obtaining richer information from the ASR decoding space, the proposed approach outperforms a strong Conformer-Transducer baseline. Another issue with standard error correction is that the generation process is not well-guided. To address this a constrained decoding process, either based on the N-best list or an ASR lattice, is used which allows additional information to be propagated.