Acoustic modeling based on model structure annealing for speech recognition
Acoustic modeling based on model structure annealing for speech recognition
复制标题
基于模型结构退火的语音识别声学建模
DOI:
10.21437/interspeech.2008-111
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
K. Tokuda
中科院分区:
文献类型:
--
作者:
Sayaka Shiota;Kei Hashimoto;H. Zen;Yoshihiko Nankaku;Akinobu Lee;K. Tokuda
This paper proposes an HMM training technique using multiple phonetic decision trees and evaluates it in speech recognition. In the use of context dependent models, the decision tree based context clustering is applied to find a parameter tying structure. However, the clustering is usually performed based on statistics of HMM state sequences which are obtained by unreliable models without context clustering. To avoid this problem, we optimize the decision trees and HMM state sequences simultaneously. In the proposed method, this is performed by maximum likelihood (ML) estimation of a newly defined statistical model which includes multiple decision trees as hidden variables. Applying the deterministic annealing expectation maximization (DAEM) algorithm and using multiple decision trees in early stage of model training, state sequences are reliably estimated. In continuous phoneme recognition experiments, the proposed method can improve the recognition performance.
DOI:
--
发表时间:
2004
期刊:
International Conference on Spoken Language Processing (INTERSPEECH2004-ICSLP2004) vol.1
影响因子:
--
作者:
Yohei Itaya;Heiga Zen;Yoshihiko Nankaku;Chiyomi Miyajima;Keiichi Tokuda;Tadashi Kitamura
通讯作者:
Tadashi Kitamura