Discriminative cluster adaptive training

Discriminative cluster adaptive training
复制标题

判别性集群自适应训练

DOI:
10.1109/tsa.2005.858555
复制
发表时间:
2006
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
M. Gales
M. Gales
中科院分区:
--
文献类型:
--
作者:
Kai Yu;M. Gales

文献摘要

被引文献

相似文献

多聚类方案,如聚类自适应训练(CAT)或特征语音系统,是一种流行的方法,用于快速扬声器和环境适应。插值权重用于将多簇、规范模型变换为代表单个扬声器或声学环境的标准隐马尔可夫模型(HMM)集。CAT的最大似然训练以前已经研究过。然而,在现有技术的大词汇量连续语音识别系统中,通常采用区分训练。本文研究了将判别式训练应用于多聚类系统。特别是,最小电话错误(MPE)更新公式的CAT系统。为了在这种情况下使用MPE,需要修改标准MPE平滑函数和与MPE训练相关联的先验分布。一个更复杂的自适应训练方案相结合的插值权重和线性变换,结构化变换(ST),也讨论了MPE训练框架内。有区别地训练CAT和ST系统进行了评估的最先进的会话电话语音任务。这些多集群系统被发现优于标准和自适应训练的系统
Multiple-cluster schemes, such as cluster adaptive training (CAT) or eigenvoice systems, are a popular approach for rapid speaker and environment adaptation. Interpolation weights are used to transform a multiple-cluster, canonical, model to a standard hidden Markov model (HMM) set representative of an individual speaker or acoustic environment. Maximum likelihood training for CAT has previously been investigated. However, in state-of-the-art large vocabulary continuous speech recognition systems, discriminative training is commonly employed. This paper investigates applying discriminative training to multiple-cluster systems. In particular, minimum phone error (MPE) update formulae for CAT systems are derived. In order to use MPE in this case, modifications to the standard MPE smoothing function and the prior distribution associated with MPE training are required. A more complex adaptive training scheme combining both interpolation weights and linear transforms, a structured transform (ST), is also discussed within the MPE training framework. Discriminatively trained CAT and ST systems were evaluated on a state-of-the-art conversational telephone speech task. These multiple-cluster systems were found to outperform both standard and adaptively trained systems