Efficient Training and Evaluation of Recurrent Neural Network Language Models for Automatic Speech Recognition

Efficient Training and Evaluation of Recurrent Neural Network Language Models for Automatic Speech Recognition
复制标题

DOI:
10.1109/taslp.2016.2598304
复制
发表时间:
2016-11-01
影响因子:
5.4
通讯作者:
Woodland, Philip C.
Woodland, Philip C.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chen, Xie;Liu, Xunying;Woodland, Philip C.

文献摘要

被引文献

相似文献

递归神经网络语言模型(RNNLM)在包括自动语音识别在内的一系列应用中越来越受欢迎。限制其可能的应用领域的一个重要问题是在训练和评估中产生的计算成本。本文描述了一系列新的效率提高方法,允许RNNLM在图形处理单元(GPU)上更有效地训练并在CPU上进行评估。首先,提出了一种基于非类的全输出层结构(F-RNNLM)的改进RNNLM架构。这种修改后的架构有助于在GPU上使用大量数据对F-RNNLM训练进行新的拼接句子束模式并行化。其次,探索了基于方差正则化和噪声对比估计的两个有效的RNNLM训练标准,以专门减少与RNNLM输出层softmax归一化项相关的计算。最后,利用多个GPU的流水线训练算法也被用来进一步提高训练速度。最初,RNNLM是在一个中等的数据集上训练的,该数据集包含来自大词汇量对话电话语音识别任务的2000万个单词。RNNLM的训练时间在单个GPU上比基于CPU的标准RNNLM工具包减少了53倍。一个56倍的速度在CPU上的测试时间评估获得了超过基线F-RNNLM。在识别精度和困惑度方面也得到了一致的改善。在谷歌10亿语料库上的实验也表明,RNNLM的训练扩展性很好。
Recurrent neural network language models (RNNLMs) are becoming increasingly popular for a range of applications including automatic speech recognition. An important issue that limits their possible application areas is the computational cost incurred in training and evaluation. This paper describes a series of new efficiency improving approaches that allows RNNLMs to be more efficiently trained on graphics processing units (GPUs) and evaluated on CPUs. First, a modified RNNLM architecture with a nonclass-based, full output layer structure (F-RNNLM) is proposed. This modified architecture facilitates a novel spliced sentence bunch mode parallelization of F-RNNLM training using large quantities of data on a GPU. Second, two efficient RNNLM training criteria based on variance regularization and noise contrastive estimation are explored to specifically reduce the computation associated with the RNNLM output layer softmax normalisation term. Finally, a pipelined training algorithm utilizing multiple GPUs is also used to further improve the training speed. Initially, RNNLMs were trained on a moderate dataset with 20M words from a large vocabulary conversational telephone speech recognition task. The training time of RNNLM is reduced by up to a factor of 53 on a single GPU over the standard CPU-based RNNLM toolkit. A 56 times speed up in test time evaluation on a CPU was obtained over the baseline F-RNNLMs. Consistent improvements in both recognition accuracy and perplexity were also obtained over C-RNNLMs. Experiments on Google's one billion corpus also reveals that the training of RNNLM scales well.