An Improved Single Step Non-autoregressive Transformer for Automatic Speech Recognition

An Improved Single Step Non-autoregressive Transformer for Automatic Speech Recognition
复制标题

DOI:
10.21437/interspeech.2021-1955
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Ruchao Fan;Wei Chu;Peng Chang;Jing Xiao;A. Alwan
Ruchao Fan;Wei Chu;Peng Chang;Jing Xiao;A. Alwan
中科院分区:
其他
文献类型:
--
作者:
Ruchao Fan;Wei Chu;Peng Chang;Jing Xiao;A. Alwan

文献摘要

相似文献

非自回归机制可以显着减少语音变换器的推理时间,特别是当单步变量被应用时。基于CTC的单步非自回归Transformer(CASS-NAT)的先前工作已经显示出比自回归transformer(AT)大的真实的时间因子(RTF)改进。在这项工作中,我们提出了几种方法来提高端到端CASS-NAT的准确性,其次是性能分析。首先,卷积增强的自注意块被应用于编码器和解码器模块。其次,我们建议扩大每个令牌的触发掩模(声学边界),以增加CTC对齐的鲁棒性。此外,迭代损失函数用于增强低层参数的梯度更新。在不使用外部语言模型的情况下,改进后的CASS-NAT在Librispeech测试clean/other集上的WER分别为3.1%/7.2%,在Aishell 1测试集上的CER为5.4%,WER/CER相对提高了7%~21%。为了进行分析,我们绘制了解码器中的注意力权重分布,以可视化标记级声学嵌入之间的关系。当声学嵌入可视化时,我们发现它们具有与词嵌入相似的行为,这解释了为什么改进的CASS-NAT的性能与AT相似。
Non-autoregressive mechanisms can significantly decrease inference time for speech transformers, especially when the single step variant is applied. Previous work on CTC alignment-based single step non-autoregressive transformer (CASS-NAT) has shown a large real time factor (RTF) improvement over autoregressive transformers (AT). In this work, we propose several methods to improve the accuracy of the end-to-end CASS-NAT, followed by performance analyses. First, convolution augmented self-attention blocks are applied to both the encoder and decoder modules. Second, we propose to expand the trigger mask (acoustic boundary) for each token to increase the robustness of CTC alignments. In addition, iterated loss functions are used to enhance the gradient update of low-layer parameters. Without using an external language model, the WERs of the improved CASS-NAT, when using the three methods, are 3.1%/7.2% on Librispeech test clean/other sets and the CER is 5.4% on the Aishell1 test set, achieving a 7%~21% relative WER/CER improvement. For the analyses, we plot attention weight distributions in the decoders to visualize the relationships between token-level acoustic embeddings. When the acoustic embeddings are visualized, we find that they have a similar behavior to word embeddings, which explains why the improved CASS-NAT performs similarly to AT.