A new phonetic tied-mixture model for efficient decoding

A new phonetic tied-mixture model for efficient decoding
复制标题

一种用于高效解码的新语音绑定混合模型

DOI:
10.1109/icassp.2000.861808
复制
发表时间:
2000
期刊:
2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.00CH37100)
影响因子:
--
通讯作者:
K. Shikano
K. Shikano
中科院分区:
--
文献类型:
--
作者:
Akinobu Lee;Tatsuya Kawahara;K. Takeda;K. Shikano

文献摘要

参考文献

被引文献

相似文献

提出了有效的大型词汇连续语音识别的语音绑定混合物(PTM)模型。它是由与上下文无关的电话模型合成的,每个状态具有64个混合物组件,通过根据Triphones的共享状态分配不同的混合物权重。然后重新估计混合物以进行优化。该模型通过20000字的报纸语料库实现了7.0%的单词错误率,这与更高分辨率的Triphone相当。与所有州共享高斯人的常规PTM相比,提出的模型易于训练并可靠地估计。此外,该模型使解码器能够进行有效的高斯修剪。据发现,计算64个组件中只有2个不会导致任何准确性损失。提出和比较了几种修剪方法,最好的方法将计算降低到约20%。
A phonetic tied-mixture (PTM) model for efficient large vocabulary continuous speech recognition is presented. It is synthesized from context-independent phone models with 64 mixture components per state by assigning different mixture weights according to the shared states of triphones. Mixtures are then re-estimated for optimization. The model achieves a word error rate of 7.0% with a 20000-word dictation of newspaper corpus, which is comparable to the best figure by the triphone of much higher resolutions. Compared with conventional PTMs that share Gaussians by all states, the proposed model is easily trained and reliably estimated. Furthermore, the model enables the decoder to perform efficient Gaussian pruning. It is found out that computing only two out of 64 components does not cause any loss of accuracy. Several methods for the pruning are proposed and compared, and the best one reduced the computation to about 20%.
T.Kawahara,T.Kobayashi,K.Takeda,N.Minematsu,K.Itou,M.Yamamoto,A.Yamada,T.Utsuro,K.Shikano:“用于日语大词汇连续语音识别的可共享软件存储库”Proc。
DOI: --
发表时间: --
期刊:
影响因子: --
作者:
通讯作者: --