Acoustic modeling based on model structure annealing for speech recognition

Acoustic modeling based on model structure annealing for speech recognition
复制标题

基于模型结构退火的语音识别声学建模

DOI:
10.21437/interspeech.2008-111
复制
发表时间:
2008
期刊:
Journal of Physics: Conference Series
影响因子:
--
通讯作者:
K. Tokuda
K. Tokuda
中科院分区:
--
文献类型:
--
作者:
Sayaka Shiota;Kei Hashimoto;H. Zen;Yoshihiko Nankaku;Akinobu Lee;K. Tokuda

文献摘要

参考文献

被引文献

相似文献

本文提出了一种基于多个语音决策树的HMM训练技术,并在语音识别中对其进行了评价。在上下文相关模型的使用中,应用基于决策树的上下文聚类来找到参数绑定结构。然而,聚类通常是基于HMM状态序列的统计数据,这些统计数据是由不可靠的模型获得的,没有上下文聚类。为了避免这个问题,我们同时优化决策树和HMM状态序列。在所提出的方法中,这是通过对新定义的统计模型进行最大似然(ML)估计来执行的,该模型包括多个决策树作为隐藏变量。应用确定性退火期望最大化(DAEM)算法,并在模型训练的早期阶段使用多个决策树,可靠地估计状态序列。在连续音素识别实验中,该方法可以提高识别性能。
This paper proposes an HMM training technique using multiple phonetic decision trees and evaluates it in speech recognition. In the use of context dependent models, the decision tree based context clustering is applied to find a parameter tying structure. However, the clustering is usually performed based on statistics of HMM state sequences which are obtained by unreliable models without context clustering. To avoid this problem, we optimize the decision trees and HMM state sequences simultaneously. In the proposed method, this is performed by maximum likelihood (ML) estimation of a newly defined statistical model which includes multiple decision trees as hidden variables. Applying the deterministic annealing expectation maximization (DAEM) algorithm and using multiple decision trees in early stage of model training, state sequences are reliably estimated. In continuous phoneme recognition experiments, the proposed method can improve the recognition performance.
声学模型参数估计中的确定性退火电磁算法
DOI: --
发表时间: 2004
期刊: International Conference on Spoken Language Processing (INTERSPEECH2004-ICSLP2004) vol.1
影响因子: --
作者:
Yohei Itaya;Heiga Zen;Yoshihiko Nankaku;Chiyomi Miyajima;Keiichi Tokuda;Tadashi Kitamura
通讯作者: Tadashi Kitamura