Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition

Context-Dependent Pre-Trained Deep Neural Networks for Large-Vocabulary Speech Recognition
复制标题

DOI:
10.1109/tasl.2011.2134090
复制
发表时间:
2012-01-01
影响因子:
--
通讯作者:
Acero, Alex
Acero, Alex
中科院分区:
其他
文献类型:
--
作者:
Dahl, George E.;Yu, Dong;Acero, Alex

文献摘要

被引文献

相似文献

我们提出了一种新的上下文相关(CD)模型的大词汇量语音识别(LVSR),利用最近的进展,使用深度信念网络的电话识别。我们描述了一种预训练的深度神经网络隐马尔可夫模型(DNN-HMM)混合架构,该架构训练DNN以产生senone(绑定三音子状态)上的分布作为其输出。深度信念网络预训练算法是一种鲁棒且通常有用的生成性初始化深度神经网络的方法,可以帮助优化并减少泛化误差。我们说明了我们的模型的关键组成部分,描述了应用CD-DNN-Hyndrome LVSR的过程,并分析了各种建模选择对性能的影响。在一个具有挑战性的商业搜索数据集上的实验表明,CD-DNN-Hacking可以显著优于传统的上下文相关高斯混合模型(GMM)-Hacking,句子准确率的绝对提高分别为5.8%和9.2%(或16.0%和23.2%的相对误差减少),分别
We propose a novel context-dependent (CD) model for large-vocabulary speech recognition (LVSR) that leverages recent advances in using deep belief networks for phone recognition. We describe a pre-trained deep neural network hidden Markov model (DNN-HMM) hybrid architecture that trains the DNN to produce a distribution over senones (tied triphone states) as its output. The deep belief network pre-training algorithm is a robust and often helpful way to initialize deep neural networks generatively that can aid in optimization and reduce generalization error. We illustrate the key components of our model, describe the procedure for applying CD-DNN-HMMs to LVSR, and analyze the effects of various modeling choices on performance. Experiments on a challenging business search dataset demonstrate that CD-DNN-HMMs can significantly outperform the conventional context-dependent Gaussian mixture model (GMM)-HMMs, with an absolute sentence accuracy improvement of 5.8% and 9.2% (or relative error reduction of 16.0% and 23.2%) over the CD-GMM-HMMs trained using the minimum phone error rate (MPE) and maximum-likelihood (ML) criteria, respectively.