Semantic Mask for Transformer based End-to-End Speech Recognition

Semantic Mask for Transformer based End-to-End Speech Recognition
复制标题

基于 Transformer 的端到端语音识别的语义掩码

DOI:
--
复制
发表时间:
2019
期刊:
Interspeech
影响因子:
--
通讯作者:
Ming Zhou
Ming Zhou
中科院分区:
--
文献类型:
--
作者:
Chengyi Wang;Yu Wu;Yujiao Du;Jinyu Li;Shujie Liu;Liang Lu;Shuo Ren;Guoli Ye;Sheng Zhao;Ming Zhou

文献摘要

被引文献

相似文献

基于注意力的编码器-解码器模型在自动语音识别(ASR)和文本转语音(TTS)任务中都取得了令人印象深刻的结果。这种方法利用神经网络的记忆能力从头开始学习从输入序列到输出序列的映射,而不需要假设先验知识(例如比对)。然而,该模型很容易出现过拟合,尤其是当训练数据量有限时。受 SpecAugment 和 BERT 的启发,在本文中,我们提出了一种基于语义掩码的正则化来训练此类端到端(E2E)模型。这个想法是屏蔽与特定输出标记(例如单词或单词片段)相对应的输入特征,以鼓励模型根据上下文信息填充标记。虽然这种方法适用于任何类型的神经网络架构的编码器-解码器框架,但我们在这项工作中研究了基于 Transformer 的 ASR 模型。我们在 Librispeech 960h 和 TedLium2 数据集上进行了实验,并在 E2E 模型范围内的测试集上实现了 state-of-the-art 的性能。
Attention-based encoder-decoder model has achieved impressive results for both automatic speech recognition (ASR) and text-to-speech (TTS) tasks. This approach takes advantage of the memorization capacity of neural networks to learn the mapping from the input sequence to the output sequence from scratch, without the assumption of prior knowledge such as the alignments. However, this model is prone to overfitting, especially when the amount of training data is limited. Inspired by SpecAugment and BERT, in this paper, we propose a semantic mask based regularization for training such kind of end-to-end (E2E) model. The idea is to mask the input features corresponding to a particular output token, e.g., a word or a word-piece, in order to encourage the model to fill the token based on the contextual information. While this approach is applicable to the encoder-decoder framework with any type of neural network architecture, we study the transformer-based model for ASR in this work. We perform experiments on Librispeech 960h and TedLium2 data sets, and achieve the state-of-the-art performance on the test set in the scope of E2E models.