Modelling of the interframe dependence in an HMM using conditional Gaussian mixtures

Modelling of the interframe dependence in an HMM using conditional Gaussian mixtures
复制标题

使用条件高斯混合对 HMM 中的帧间依赖性进行建模

DOI:
10.1006/csla.1996.0012
复制
发表时间:
1996
期刊:
Comput. Speech Lang.
影响因子:
--
通讯作者:
F. J. Smith
F. J. Smith
中科院分区:
--
文献类型:
--
作者:
J. Ming;F. J. Smith

文献摘要

被引文献

相似文献

摘要 本文研究了语音识别隐马尔可夫模型(HMM)中帧间依赖性的建模。首先,提出了一种新的观察模型,假设依赖于多个先前帧。该模型表示这样一种依赖结构,其中包含一组一阶条件高斯密度的加权混合,每个混合分量代表一个特定的条件框架。接下来,在训练和识别中执行条件帧/片段选择的优化,从而有助于消除由于不同观察历史而导致的条件片段的不匹配。开发了 EM(期望最大化)迭代算法,用于估计模型参数和优化依赖结构。在与说话人无关的 E-set 数据库上进行的实验比较表明,在没有对依赖结构进行优化的情况下,新模型在参数大小相当或更小的情况下,比标准 HMM、二元组 HMM 和线性预测 HMM 获得了更好的性能。对依赖结构的优化导致性能的进一步提升。
Abstract This paper investigates the modelling of the interframe dependence in a hidden Markov model (HMM) for speech recognition. First, a new observation model, assuming dependence on multiple previous frames, is proposed. This model represents such a dependence structure with a weighted mixture of a set of first-order conditional Gaussian densities, each mixture component accounting for a specific conditional frame. Next, an optimization in choosing the conditional frames/segment is performed in both training and recognition, thereby helping to remove the mismatch of the conditional segments due to different observation histories. An EM (Expectation–Maximization) iteration algorithm is developed for the estimation of the model parameters and for the optimization over the dependence structure. Experimental comparisons on a speaker-independent E-set database show that the new model, without optimization on the dependence structure, achieves better performance than the standard HMM, the bigram HMM and the linear-predictive HMM, all in comparable or smaller parameter sizes. The optimization over the dependence structure leads to further improvement in the performance.