Fast End-to-End Speech Recognition Via Non-Autoregressive Models and Cross-Modal Knowledge Transferring From BERT

Fast End-to-End Speech Recognition Via Non-Autoregressive Models and Cross-Modal Knowledge Transferring From BERT
复制标题

通过非自回归模型和 BERT 的跨模态知识传输进行快速端到端语音识别

DOI:
10.1109/taslp.2021.3082299
复制
发表时间:
2021-01-01
影响因子:
5.4
通讯作者:
Zhang, Shuai
Zhang, Shuai
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bai, Ye;Yi, Jiangyan;Zhang, Shuai

文献摘要

被引文献

相似文献

基于注意力的编码器-解码器(AED)模型在语音识别中取得了令人满意的性能。然而,由于解码器以自回归方式预测文本标记(例如字符或单词),因此AED模型难以并行预测所有标记。这使得推理速度相对较慢。相比之下,我们提出了一个端到端的非自回归语音识别模型,称为LASO(认真倾听,拼写一次)。该模型通过注意机制将编码的语音特征聚集到与每个标记对应的隐藏表示中。因此,该模型可以捕获令牌关系的自注意力的聚合隐藏表示从整个语音信号,而不是自回归模型的令牌。在没有显式自回归语言建模的情况下,该模型并行预测序列中的所有标记,从而使推理效率更高。此外,我们提出了一种跨模态迁移学习方法,使用文本模态语言模型,通过对齐令牌语义来提高语音模态LASO的性能。我们在两个尺度的公开汉语语音数据集AISHELL-1和AISHELL-2上进行了实验。实验结果表明,与自回归Transformer模型相比,该模型具有约50倍的加速比和较好的性能.文本模态模型的跨模态知识传递可以提高语音模态模型的性能。
Attention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because the decoder predicts text tokens (such as characters or words) in an autoregressive manner, it is difficult for an AED model to predict all tokens in parallel. This makes the inference speed relatively slow. In contrast, we propose an end-to-end non-autoregressive speech recognition model called LASO (Listen Attentively, and Spell Once). The model aggregates encoded speech features into the hidden representations corresponding to each token with attention mechanisms. Thus, the model can capture the token relations by self-attention on the aggregated hidden representations from the whole speech signal rather than autoregressive modeling on tokens. Without explicitly autoregressive language modeling, this model predicts all tokens in the sequence in parallel so that the inference is efficient. Moreover, we propose a cross-modal transfer learning method to use a text-modal language model to improve the performance of speech-modal LASO by aligning token semantics. We conduct experiments on two scales of public Chinese speech datasets AISHELL-1 and AISHELL-2. Experimental results show that our proposed model achieves a speedup of about $50\times$ and competitive performance, compared with the autoregressive transformer models. And the cross-modal knowledge transferring from the text-modal model can improve the performance of the speech-modal model.