Improved Mask-CTC for Non-Autoregressive End-to-End ASR

Improved Mask-CTC for Non-Autoregressive End-to-End ASR
复制标题

DOI:
10.1109/icassp39728.2021.9414198
复制
发表时间:
2020-10
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Yosuke Higuchi;H. Inaguma;Shinji Watanabe;Tetsuji Ogawa;Tetsunori Kobayashi
Yosuke Higuchi;H. Inaguma;Shinji Watanabe;Tetsuji Ogawa;Tetsunori Kobayashi
中科院分区:
其他
文献类型:
--
作者:
Yosuke Higuchi;H. Inaguma;Shinji Watanabe;Tetsuji Ogawa;Tetsunori Kobayashi

文献摘要

被引文献

相似文献

对于自动语音识别(ASR)的实际部署,期望系统能够快速推理,同时减轻对计算资源的需求。最近提出的端到端ASR系统的基础上掩码预测与连接主义时间分类(CTC),掩码CTC,满足这一需求,通过生成令牌在一个非自回归的方式。虽然Mask-CTC实现了非常快的推理速度,但其识别性能福尔斯落后于传统的自回归(AR)系统。为了提高Mask-CTC的性能,我们首先建议通过采用最近提出的称为Conformer的架构来增强编码器网络架构。接下来,我们提出了新的训练和解码方法,通过引入辅助目标来预测部分目标序列的长度,这允许模型在推理过程中删除或插入令牌。不同ASR任务的实验结果表明,所提出的方法显着提高了Mask-CTC,优于标准CTC模型(WSJ上的15.5% → 9.1% WER)。此外,Mask-CTC现在实现了与AR模型竞争的结果,而不会降低推理速度(使用CPU时< 0.1 RTF)。我们还展示了Mask-CTC在端到端语音翻译中的潜在应用。
For real-world deployment of automatic speech recognition (ASR), the system is desired to be capable of fast inference while relieving the requirement of computational resources. The recently proposed end-to-end ASR system based on mask-predict with connectionist temporal classification (CTC), Mask-CTC, fulfills this demand by generating tokens in a non-autoregressive fashion. While Mask-CTC achieves remarkably fast inference speed, its recognition performance falls behind that of conventional autoregressive (AR) systems. To boost the performance of Mask-CTC, we first propose to enhance the encoder network architecture by employing a recently proposed architecture called Conformer. Next, we propose new training and decoding methods by introducing auxiliary objective to predict the length of a partial target sequence, which allows the model to delete or insert tokens during inference. Experimental results on different ASR tasks show that the proposed approaches improve Mask-CTC significantly, outperforming a standard CTC model (15.5% → 9.1% WER on WSJ). Moreover, Mask-CTC now achieves competitive results to AR models with no degradation of inference speed (< 0.1 RTF using CPU). We also show a potential application of Mask-CTC to end-to-end speech translation.