Modelling of the interframe dependence in an HMM using conditional Gaussian mixtures
Modelling of the interframe dependence in an HMM using conditional Gaussian mixtures
复制标题
使用条件高斯混合对 HMM 中的帧间依赖性进行建模
DOI:
10.1006/csla.1996.0012
复制
发表时间:
1996
期刊:
影响因子:
--
通讯作者:
F. J. Smith
中科院分区:
文献类型:
--
作者:
J. Ming;F. J. Smith
Abstract This paper investigates the modelling of the interframe dependence in a hidden Markov model (HMM) for speech recognition. First, a new observation model, assuming dependence on multiple previous frames, is proposed. This model represents such a dependence structure with a weighted mixture of a set of first-order conditional Gaussian densities, each mixture component accounting for a specific conditional frame. Next, an optimization in choosing the conditional frames/segment is performed in both training and recognition, thereby helping to remove the mismatch of the conditional segments due to different observation histories. An EM (Expectation–Maximization) iteration algorithm is developed for the estimation of the model parameters and for the optimization over the dependence structure. Experimental comparisons on a speaker-independent E-set database show that the new model, without optimization on the dependence structure, achieves better performance than the standard HMM, the bigram HMM and the linear-predictive HMM, all in comparable or smaller parameter sizes. The optimization over the dependence structure leads to further improvement in the performance.