Improving Non-Autoregressive End-to-End Speech Recognition with Pre-Trained Acoustic and Language Models

Improving Non-Autoregressive End-to-End Speech Recognition with Pre-Trained Acoustic and Language Models
复制标题

DOI:
10.1109/icassp43922.2022.9746316
复制
发表时间:
2022-01
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Keqi Deng;Zehui Yang;Shinji Watanabe;Yosuke Higuchi;Gaofeng Cheng;Pengyuan Zhang
Keqi Deng;Zehui Yang;Shinji Watanabe;Yosuke Higuchi;Gaofeng Cheng;Pengyuan Zhang
中科院分区:
其他
文献类型:
--
作者:
Keqi Deng;Zehui Yang;Shinji Watanabe;Yosuke Higuchi;Gaofeng Cheng;Pengyuan Zhang

文献摘要

相似文献

虽然变形金刚在端到端(E2E)自动语音识别(ASR)方面取得了可喜的成果,但其自回归(AR)结构成为加快解码过程的瓶颈。对于现实世界的部署,ASR系统要求在实现快速推理的同时具有很高的精度。非自回归(NAR)模型因其快速的推理速度而成为一种流行的选择,但在识别精度方面仍落后于AR系统。为了满足这两个需求,本文提出了一种NAR CTC/注意模型,该模型利用预先训练的声学模型和语言模型:Wave2ve2.0和BERT。为了弥合从预先训练的模型获得的语音和文本表示之间的情态差异,我们设计了一种新的情态转换机制,该机制更适合于标志语言。在推理过程中,我们使用CTC分支来生成目标长度,这使得BERT能够并行地预测令牌。我们还设计了一种基于缓存的CTC/注意联合译码方法,在保持译码速度较快的同时提高了识别精度。实验结果表明,所提出的NAR模型的性能大大优于我们的Wave2ve2.0 CTC基线(相对于AISHELL-1的相对CER降低了15.1%)。所提出的NAR模型在AISHELL-1基准上显著超过了以前的NAR系统,并显示了在英语任务中的潜力。
While Transformers have achieved promising results in end-to-end (E2E) automatic speech recognition (ASR), their autoregressive (AR) structure becomes a bottleneck for speeding up the decoding process. For real-world deployment, ASR systems are desired to be highly accurate while achieving fast inference. Non-autoregressive (NAR) models have become a popular alternative due to their fast inference speed, but they still fall behind AR systems in recognition accuracy. To fulfill the two demands, in this paper, we propose a NAR CTC/attention model utilizing both pre-trained acoustic and language models: wav2vec2.0 and BERT. To bridge the modality gap between speech and text representations obtained from the pre-trained models, we design a novel modality conversion mechanism, which is more suitable for logographic languages. During inference, we employ a CTC branch to generate a target length, which enables the BERT to predict tokens in parallel. We also design a cache-based CTC/attention joint decoding method to improve the recognition accuracy while keeping the decoding speed fast. Experimental results show that the proposed NAR model greatly outperforms our strong wav2vec2.0 CTC baseline (15.1% relative CER reduction on AISHELL-1). The proposed NAR model significantly surpasses previous NAR systems on the AISHELL-1 benchmark and shows a potential for English tasks.