Deep Neural Networks for Acoustic Modeling in Speech Recognition

Deep Neural Networks for Acoustic Modeling in Speech Recognition
复制标题

DOI:
10.1109/msp.2012.2205597
复制
发表时间:
2012-11-01
影响因子:
14.9
通讯作者:
Kingsbury, Brian
Kingsbury, Brian
中科院分区:
工程技术1区
文献类型:
--
作者:
Hinton, Geoffrey;Deng, Li;Kingsbury, Brian

文献摘要

被引文献

相似文献

大多数当前的语音识别系统使用隐马尔可夫模型(HMM)来处理语音的时间变化性,并使用高斯混合模型(GMM)来确定每个HMM的每个状态与代表声学输入的系数帧或系数帧短窗口的拟合程度。评估拟合的另一种方法是使用前馈神经网络,该网络将几帧系数作为输入,并产生HMM状态的后验概率作为输出。具有许多隐藏层并使用新方法进行训练的深度神经网络(DNN)已被证明在各种语音识别基准测试中优于Gandhi,有时甚至是大幅度的。本文概述了这一进展,并代表了最近成功使用DNN进行语音识别声学建模的四个研究小组的共同观点。
Most current speech recognition systems use hidden Markov models (HMMs) to deal with the temporal variability of speech and Gaussian mixture models (GMMs) to determine how well each state of each HMM fits a frame or a short window of frames of coefficients that represents the acoustic input. An alternative way to evaluate the fit is to use a feed-forward neural network that takes several frames of coefficients as input and produces posterior probabilities over HMM states as output. Deep neural networks (DNNs) that have many hidden layers and are trained using new methods have been shown to outperform GMMs on a variety of speech recognition benchmarks, sometimes by a large margin. This article provides an overview of this progress and represents the shared views of four research groups that have had recent successes in using DNNs for acoustic modeling in speech recognition.