Eigenvalues Driven Gaussian Selection in continuous speech recognition using HMMs with full covariance matrices

Eigenvalues Driven Gaussian Selection in continuous speech recognition using HMMs with full covariance matrices
复制标题

DOI:
10.1007/s10489-008-0152-9
复制
发表时间:
2010-10
影响因子:
5.3
通讯作者:
Marko Janev;D. Pekar;N. Jakovljević;V. Delić
Marko Janev;D. Pekar;N. Jakovljević;V. Delić
中科院分区:
计算机科学2区
文献类型:
--
作者:
Marko Janev;D. Pekar;N. Jakovljević;V. Delić

文献摘要

被引文献

相似文献

本文提出了一种用于连续语音识别系统的混合语音高斯选择算法。该系统是基于隐马尔可夫模型(HMM),使用高斯混合与全协方差矩阵作为输出分布。高斯选择的目的是提高语音识别系统的速度,而不降低识别精度。其基本思想是通过矢量量化(VQ)将接近的混合物聚类成单个组,并为其分配唯一的高斯参数进行估计,从而形成超混合物。在解码过程中,仅选择那些高于指定阈值的超混合物,并且仅评估属于它们的混合物,从而提高计算效率。如果混合物之间的重叠很小,并且它们的方差在同一范围内,则聚类和评估没有问题。然而,在真实的情况下,有许多模型不符合此配置文件。本文提出的高斯选择方案解决了这个问题。为此,除了聚类算法之外,它还结合了混合分组算法。基于使用有序加权平均算子(OWA)从该混合的协方差矩阵的特征值聚合的值,将特定混合分配给来自预定义组集合的组。在对混合物进行分组后,分别对每组进行高斯混合聚类。
In this paper a novel algorithm for Gaussian Selection (GS) of mixtures used in a continuous speech recognition system is presented. The system is based on hidden Markov models (HMM), using Gaussian mixtures with full covariance matrices as output distributions. The purpose of Gaussian selection is to increase the speed of a speech recognition system, without degrading the recognition accuracy. The basic idea is to form hyper-mixtures by clustering close mixtures into a single group by means of Vector Quantization (VQ) and assigning it unique Gaussian parameters for estimation. In the decoding process only those hyper-mixtures which are above a designated threshold are selected, and only mixtures belonging to them are evaluated, improving computational efficiency. There is no problem with the clustering and evaluation if overlaps between the mixtures are small, and their variances are of the same range. However, in real case, there are numerous models which do not fit this profile. A Gaussian selection scheme proposed in this paper addresses this problem. For that purpose, beside the clustering algorithm, it also incorporates an algorithm for mixture grouping. The particular mixture is assigned to a group from the predefined set of groups, based on a value aggregated from eigenvalues of the covariance matrix of that mixture using Ordered Weighted Averaging operators (OWA). After the grouping of mixtures is carried out, Gaussian mixture clustering is performed on each group separately.