A Study of Variable-Parameter Gaussian Mixture Hidden Markov Modeling for Noisy Speech Recognition

A Study of Variable-Parameter Gaussian Mixture Hidden Markov Modeling for Noisy Speech Recognition
复制标题

DOI:
10.1109/tasl.2006.889791
复制
发表时间:
2007-05
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Xiaodong Cui;Y. Gong
Xiaodong Cui;Y. Gong
中科院分区:
其他
文献类型:
--
作者:
Xiaodong Cui;Y. Gong

文献摘要

被引文献

相似文献

为了提高噪声环境中的识别性能,通常采用多条件训练,其中将被各种噪声破坏的语音信号用于声学模型训练。已发表的语音隐马尔可夫建模使用多个高斯分布来覆盖噪声引起的语音分布的扩散,这分散了语音事件本身的建模,并可能牺牲干净语音的性能。在本文中,我们提出了一种新方法,通过将状态发射参数(均值和方差)建模为连续环境相关变量的多项式函数来扩展传统高斯混合隐马尔可夫模型(GMHMM)。在识别时,一组特定于环境变量给定值的HMM被实例化并用于识别。所提出的可变参数 GMHMM 多项式函数的最大似然 (ML) 估计在期望最大化 (EM) 框架内给出。 Aurora 2数据库上的实验表明,与传统的GMHMM相比,变参数高斯混合HMM有显着的改进
To improve recognition performance in noisy environments, multicondition training is usually applied in which speech signals corrupted by a variety of noise are used in acoustic model training. Published hidden Markov modeling of speech uses multiple Gaussian distributions to cover the spread of the speech distribution caused by noise, which distracts the modeling of speech event itself and possibly sacrifices the performance on clean speech. In this paper, we propose a novel approach which extends the conventional Gaussian mixture hidden Markov model (GMHMM) by modeling state emission parameters (mean and variance) as a polynomial function of a continuous environment-dependent variable. At the recognition time, a set of HMMs specific to the given value of the environment variable is instantiated and used for recognition. The maximum-likelihood (ML) estimation of the polynomial functions of the proposed variable-parameter GMHMM is given within the expectation-maximization (EM) framework. Experiments on the Aurora 2 database show significant improvements of the variable-parameter Gaussian mixture HMMs compared to the conventional GMHMMs