Acoustic modeling based on the MDL principle for speech recognition

Acoustic modeling based on the MDL principle for speech recognition
复制标题

基于MDL原理的语音识别声学建模

DOI:
10.21437/eurospeech.1997-52
复制
发表时间:
1997
期刊:
5th International Conference on Spoken Language Processing (ICSLP 1998)
影响因子:
--
通讯作者:
Takao Watanabe
Takao Watanabe
中科院分区:
--
文献类型:
--
作者:
K. Shinoda;Takao Watanabe

文献摘要

被引文献

相似文献

基于MDL原理的语音识别声学建模KoichiShinoda和TakaoWatanabeNECCorp oration 4 -1-1宫崎,Miyamae-ku,川崎216,日本fshinoda,eg@hum.cl摘要最近,上下文相关的电话单元,如三电话,B已被用于在基于隐马尔可夫模型(HMM)的语音识别中对子字单元进行建模。虽然大多数此类方法采用HMM参数的聚类(例如,在一个实施例中,子字聚类、状态聚类等),为了控制隐马尔可夫模型的大小,以避免由于插入而导致的定位或识别准确性,训练数据的可信度很低,没有一个能为B应该执行的最佳聚类程度提供有效的标准。本文提出了一种用语音决策树实现状态聚类的方法,并利用MDL准则优化聚类度。大量的日语语音识别实验表明,在用传统的聚类方法得到的各种大小的语音识别模型中,用这种方法得到的模型的识别精度最高。使用与上下文相关的(CD)电话模型而不是与上下文相关的(CI)电话模型(单声道),可以提高识别准确性[1-7]。由于CD模型的数量通常比CI模型的数量大得多,因此使用CD模型可以更好地捕捉语音数据中的变化。B字母。然而,可用训练数据的不足可能会导致B e insult。有能力支持如此大量的CD模型的使用。准备这么大量的数据往往是不切实际的。此外,CD phone应用程序收听训练数据的频率通常在CD phone集合中基本上不同;在大多数情况下,某些CD phone的频率非常小,以至于即使提供大量数据,这些CD phone也不会收听训练数据。此数据插入科学往往造成语音识别性能的严重退化。大多数使用CD模型的识别系统都采用模型参数的聚类来缓解部分问题,为此B已经发展了各种各样的聚类方法艾德。首先,对进行聚类的单元有几种选择; Leeet al. [1]例如,使用子字聚类,Hwanget al. [2]使用stateclustering,和Digalakis et al. [3]首先,提出了一种基于高斯混合态密度的混合成分聚类方法,其次,提出了几种选择声学相似单元进行B聚类的方法,有些方法只利用数据的声学特征,以自底向上的方式进行单元的合并[4,2,3 ]。另外,其他方法利用关于单元之间的声学相似性的优先级知识,其主要由决策树表示[1,5,6,7]。在后一种方法中,CI模型单元的分裂是自顶向下进行的,而不是CD模型单元的合并,在这些聚类方法中,重要的是正确地度量单元化训练数据之间的声学相似性B,最成功的方法之一是基于最大似然(ML)准则的方法(例如,在下文中,为了简单起见,解释了分裂方法(自顶向下聚类),尽管类似的解释也适用于合并方法(B自顶向上聚类)。在这种方法中,分裂的可能性的增加是为每个单元集合中的单元计算的,并且具有最大增加的单元被选择和分裂。在分裂的最后阶段,模型集B与不聚类的CD模型集几乎相同。因此,这种方法需要一个外部参数来控制聚类的程度,大多数方法使用一个阈值来限制分裂,该阈值取决于可能性的增加或单元的数目。这些阈值需要通过一系列使用测试数据的验证实验或交叉验证方法进行优化B。这些优化过程计算量大,需要的数据量大,而且没有很强的理论依据.本文提出了一种新的聚类方法,即用最小描述长度(MDL)准则代替ML准则. MDL方法[9 ]是基于信息论准则的,其B已经用于为给定的数据集选择具有适当复杂度的概率模型。该MDL准则不仅对选择B拆分单元有效,而且对决定是否停止拆分也有效。因此,不需要其他外部参数来控制聚类程度。MDL CRITERIONMDL[9]是一个信息准则,它在从众多的概率模型中选择最优模型时有B甚至被证明是B有效的。MDL准则从一组模型中选择对给定数据具有最小描述长度的模型作为最优模型。当给定一组模型f 1;:;i;:; Ig时,数据f=1;:;xNg的描述长度li(xN)以及基础模型由下式给出:(i)xN)+i2 N + logI(1)其中i是模型i的维数(自由参数的数目),(i)is参数的最大似然估计(i)=(1;:;i)模式i.(1)中的第一项是当模型被用作概率模型时,dataxN的代码长度。该术语
ACOUSTIC MODELING BASED ON THE MDLPRINCIPLE FOR SPEECH RECOGNITIONKoichi Shinoda and Takao WatanabeNEC Corp oration4-1-1 Miyazaki, Miyamae-ku, Kawasaki 216, JAPANfshino da,watanab eg@hum.cl .nec.co.jpABSTRACTRecently context-dep endent phone units, such as tri-phones, have b een used to mo del subword units in sp eechrecognition based on Hidden Markov Mo dels (HMMs).While most such metho ds employ clustering of theHMM parameters(e.g., subword clustering, state cluster-ing, etc.), to control HMM size so as to avoid p o or recogni-tion accuracy due to an insuciency of training data, noneof them provide any e ective criterion for the optimal de-gree of clustering that should b e p erformed. This pap erprop oses a metho d in which state clustering is accom-plished byway of phonetic decision trees and in which theMDL criterion is used to optimize the degree of cluster-ing. Large-vo cabulary Japanese recognition exp erimentsshow that the mo dels obtained by this metho d achievedthe highest accuracy among the mo dels of various sizesobtained with conventional clustering approaches.1.INTRODUCTIONOver the past few years, extensive studies have b een car-ried out on sp eaker-indep end ent sp eech recognition us-ing continuous density Hidden Markov Mo dels (HMMs).It is well known that in most such systems, the use ofcontext-dep endent(CD) phone mo dels instead of context-indep endent(CI) phone mo dels(monophon es), improvesrecognition accuracy[1-7].Since the numb er of CD mo dels is usually much largerthan that of CI mo dels, using CD mo dels b etter capturesvariations in sp eech data. However, the amountof aail-able training data is likely to b e insucient to supp ortthe use of such a large numb er of CD mo dels. It is oftenimpractical to prepare such a large amount of data. Fur-thermore, the frequency with which a CD phone app earsin training data usually di ers substantiall y in the set ofCD phones; in most case, the frequencies for some CDphones are so small that those CD phones do not app earin training data even if a large amount of data is pro-vided. This data insuciency often causes serious degra-dation in sp eech recognition p erformance. Most recogni-tion systems using CD mo dels employ clustering of mo delparameters to try to alleviate part of the problem.Various clustering metho ds have b een develop ed for thispurp ose. First, there are several choices for the units towhich clustering is carried out; K.F. Leeet al.[1], for ex-ample, use subword clustering, Hwanget al.[2] use stateclustering, and Digalakis et al.[3] cluster the mixture com-p onents of the HMMs with Gaussian-mixture state ob-servation densities.Second, there are several metho dsto select the acoustically-si mi lar units to b e clustered.Some metho ds use only the acoustic characteristics of thedata and the merging of the units are carried out in ab ottom-up manner[4 , 2, 3 ]. The other metho ds, in addi-tion, utilizea prioriknowledge ab out acoustic similariti esbetween the units, which are mostly represented by deci-sion trees[1, 5, 6, 7]. In most of the latter metho ds, split-ting of the units of CI mo dels is carried out in a top-downmanner, instead of merging the units of CD mo dels.In these clustering metho ds, it is imp ortant to prop-erly measure the acoustic similariti es b etween the unitsutilizing training data, in order to select the units tob e clustered.One of the most successful approachis the approach based on the maximum-likel ihood(ML)criterion(e.g.,[7 ]).In the following,for simplicity,the splitting metho d(top-down clustering) is explained,though the similar explanation is also applicable to themerging metho d(b ottom-up clustering). In this approach,the increase of the likeliho o d by splitting is calculated foreach unit in the unit set, and the unit that has the largestincrease is selected and split.However, this ML approach has one drawback. In mostcase, the likelihood becomes larger as the numb er of unitsb ecomes larger.In the nal stage of the splitting, themo del set b ecomes almost identical to the set of CD mo d-els without clustering. Therefore, this approach requiresan external parameter to control the degree of clustering.Most metho ds limit splitting using a threshold on the in-crease in the likeliho o d or on the numb er of units. Thesethresholds needs to b e optimized through a series of recog-nition exp eriments using test data or by a cross-validationmetho d. These optimization pro cesses are computation-ally exp ensive, need more data, and have no strong theo-retical justi cation.In this pap er we prop ose a new approach in whicha minimum description length(MDL) criterion, insteadof the ML criterion, is used for clustering.The MDLapproach[9 ] is based on an information theoretic criterion,which has b een used for selecting the probabilisti c mo delwith an appropriate complexity for the given amountofdata. This MDL criterion is e ective not only for select-ing the units to b e split, but also for deciding whether tostop splitting. Therefore, no other external parameter isneeded to control the degree of clustering. We apply thiscriterion to state splitting using phonetic decision tree.2.MDL CRITERIONMDL[9] is an information criterion which has b een provento b e e ective in selecting the optimal mo del from amongvarious probabilis tic mo dels. The MDL criterion selectsthe mo del with the minimum description length for thegiven data as the optimal mo del from among a set of mo d-els. When a set of mo delsf1; :::;i;:::;Igis given, the de-scription length,li(xN), of the data,f=1;:::;xNg,together with an underlying mo deliis given by,l(i)=logP^(i)xN)+i2N+ logI(1)whereiis the dimensionali ty (the numb er of free param-eters) of mo deli, and^(i)is the maximum likeliho o d es-timates for the parameters(i)=(1;:::;i)ofmodeli. The rst term in (1) is the co de length for the dataxNwhen mo deliis used as a probabili stic mo del. This term