Convolutional Neural Networks for Speech Recognition

Convolutional Neural Networks for Speech Recognition
复制标题

DOI:
10.1109/taslp.2014.2339736
复制
发表时间:
2014-10-01
影响因子:
5.4
通讯作者:
Yu, Dong
Yu, Dong
中科院分区:
计算机科学2区
文献类型:
--
作者:
Abdel-Hamid, Ossama;Mohamed, Abdel-Rahman;Yu, Dong

文献摘要

被引文献

相似文献

最近,混合深度神经网络(DNN) - 隐马尔可夫模型(HMM)已被证明在语音识别性能上相较于传统的高斯混合模型(GMM) - HMM有显著提高。性能的提升部分归因于DNN对语音特征中复杂相关性进行建模的能力。在本文中,我们表明通过使用卷积神经网络(CNNs)可以进一步降低错误率。我们首先简要描述基本的CNN,并解释它如何可用于语音识别。我们进一步提出一种有限权重共享方案,它能够更好地对语音特征进行建模。CNN中的局部连接、权重共享和池化等特殊结构对语音特征沿频率轴的小偏移表现出一定程度的不变性,这对于处理说话人和环境的变化非常重要。实验结果表明,在TIMIT音素识别和语音搜索大词汇量语音识别任务中,与DNNs相比,CNNs将错误率降低了6% - 10%。
Recently, the hybrid deep neural network (DNN)hidden Markov model (HMM) has been shown to significantly improve speech recognition performance over the conventional Gaussian mixture model (GMM)-HMM. The performance improvement is partially attributed to the ability of the DNN to model complex correlations in speech features. In this paper, we show that further error rate reduction can be obtained by using convolutional neural networks (CNNs). We first present a concise description of the basic CNN and explain how it can be used for speech recognition. We further propose a limited-weight-sharing scheme that can better model speech features. The special structure such as local connectivity, weight sharing, and pooling in CNNs exhibits some degree of invariance to small shifts of speech features along the frequency axis, which is important to deal with speaker and environment variations. Experimental results show that CNNs reduce the error rate by 6%-10% compared with DNNs on the TIMIT phone recognition and the voice search large vocabulary speech recognition tasks.