Integrating Knowledge Into End-to-End Speech Recognition From External Text-Only Data

Integrating Knowledge Into End-to-End Speech Recognition From External Text-Only Data
复制标题

将来自外部纯文本数据的知识集成到端到端语音识别中

DOI:
10.1109/taslp.2021.3066274
复制
发表时间:
2021
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Shuai Zhang
Shuai Zhang
中科院分区:
其他
文献类型:
--
作者:
Ye Bai;Jiangyan Yi;Jianhua Tao;Zhengqi Wen;Zhengkun Tian;Shuai Zhang

文献摘要

相似文献

基于注意力的编码器-解码器(AED)模型在语音识别中取得了很好的效果。然而,由于是端到端的训练,AED模型通常使用语音-文本配对数据进行训练。将外部纯文本数据合并到AED模型中是一项挑战。AED模型的另一个问题是,它在预测令牌时没有使用文本令牌的正确上下文。为了缓解上述两个问题,我们提出了一种名为LST (Learn Spelling from Teachers)的统一方法,将外部纯文本数据中的知识整合到AED模型中,并在句子中利用整个上下文。该方法分为两个阶段。首先,在表征阶段,对文本进行语言模型训练。可以看出,文本中的知识被压缩到LM中。然后,在转移阶段,通过师生学习将知识转移到AED模型中。为了进一步使用文本句子的整个上下文,我们提出了一个称为因果完形补全器(COR)的LM,它在给定左上下文和右上下文的情况下估计一个令牌的概率。因此,通过LST训练,AED模型可以利用句子中的整个上下文。与基于融合的方法在解码过程中使用LM不同,该方法在推理阶段不会增加任何额外的复杂性。我们在两个尺度的中文公开数据集AISHELL-1和AISHELL-2上进行了实验。实验结果表明,与基线混合系统和基于AED模型的系统相比,本文提出的方法可以有效地利用外部纯文本数据和句子中的整个上下文。
Attention-based encoder-decoder (AED) models have achieved promising performance in speech recognition. However, because of the end-to-end training, an AED model is usually trained with speech-text paired data. It is challenging to incorporate external text-only data into AED models. Another issue of the AED model is that it does not use the right context of a text token while predicting the token. To alleviate the above two issues, we propose a unified method called LST (Learn Spelling from Teachers) to integrate knowledge into an AED model from the external text-only data and leverage the whole context in a sentence. The method is divided into two stages. First, in the representation stage, a language model is trained on the text. It can be seen as that the knowledge in the text is compressed into the LM. Then, at the transferring stage, the knowledge is transferred to the AED model via teacher-student learning. To further use the whole context of the text sentence, we propose an LM called causal cloze completer (COR), which estimates the probability of a token, given both the left context and the right context of it. Therefore, with LST training, the AED model can leverage the whole context in the sentence. Different from fusion based methods, which use LM during decoding, the proposed method does not increase any extra complexity at the inference stage. We conduct experiments on two scales of public Chinese datasets AISHELL-1 and AISHELL-2. The experimental results demonstrate the effectiveness of leveraging external text-only data and the whole context in a sentence with our proposed method, compared with baseline hybrid systems and AED model based systems.