Language Model Integration Based on Memory Control for Sequence to Sequence Speech Recognition

Language Model Integration Based on Memory Control for Sequence to Sequence Speech Recognition
复制标题

DOI:
10.1109/icassp.2019.8683380
复制
发表时间:
2018-11
期刊:
ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Jaejin Cho;Shinji Watanabe;Takaaki Hori;M. Baskar;H. Inaguma;J. Villalba;N. Dehak
Jaejin Cho;Shinji Watanabe;Takaaki Hori;M. Baskar;H. Inaguma;J. Villalba;N. Dehak
中科院分区:
其他
文献类型:
--
作者:
Jaejin Cho;Shinji Watanabe;Takaaki Hori;M. Baskar;H. Inaguma;J. Villalba;N. Dehak

文献摘要

相似文献

在本文中,我们探索了几种新的方案来训练seq2seq模型,以集成预训练的语言模型(LM)。我们提出的融合方法专注于seq2seq解码器长短期记忆(LSTM)中的记忆单元状态和隐藏状态,与以往的研究不同,记忆单元状态由LM更新。这意味着主seq2seq保留的内存将由外部LM调整。这些融合方法有几个变体,这取决于该存储单元更新的架构以及直接影响最终标签推断的存储单元和隐藏状态的使用。我们进行了实验,以显示所提出的方法的有效性,在一个单语言的ASR设置上的Librispeech语料库,并在一个转移学习设置从多语言的ASR(MLASR)的基础模型,以低资源的语言。在Librispeech中,我们的最佳模型相对于浅融合基线提高了3.7%的WER,对于测试干净,测试其他,提高了2.4%,具有多级解码。在从MLASR基础模型到IARPA Babel Swahili模型的迁移学习中,相对于2阶段迁移基线,最佳方案将eval集上的迁移模型提高了9.9%,在CER、WER中提高了9.8%。
In this paper, we explore several new schemes to train a seq2seq model to integrate a pre-trained language model (LM). Our proposed fusion methods focus on the memory cell state and the hidden state in the seq2seq decoder long short-term memory (LSTM), and the memory cell state is updated by the LM unlike the prior studies. This means the memory retained by the main seq2seq would be adjusted by the external LM. These fusion methods have several variants depending on the architecture of this memory cell update and the use of memory cell and hidden states which directly affects the final label inference. We performed the experiments to show the effectiveness of the proposed methods in a mono-lingual ASR setup on the Librispeech corpus and in a transfer learning setup from a multilingual ASR (MLASR) base model to a low-resourced language. In Librispeech, our best model improved WER by 3.7%, 2.4% for test clean, test other relatively to the shallow fusion baseline, with multilevel decoding. In transfer learning from an MLASR base model to the IARPA Babel Swahili model, the best scheme improved the transferred model on eval set by 9.9%, 9.8% in CER, WER relatively to the 2-stage transfer baseline.