Automatic Complexity Control of Generalized Variable Parameter HMMs for Noise Robust Speech Recognition

Automatic Complexity Control of Generalized Variable Parameter HMMs for Noise Robust Speech Recognition
复制标题

用于噪声鲁棒语音识别的广义变参数 HMM 的自动复杂度控制

DOI:
10.1109/taslp.2014.2372901
复制
发表时间:
2015
影响因子:
5.4
通讯作者:
Lan Wang
Lan Wang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Rongfeng Su;Xunying Liu;Lan Wang

文献摘要

参考文献

相似文献

自动语音识别(ASR)系统的声学建模问题的一个重要部分是处理与由时变外部因素(例如环境噪声)创建的目标环境的失配。这个问题的一个可能的解决方案是引入可控性的基础声学模型,以允许瞬时适应的基础噪声条件。沿着这条线,可以使用例如广义可变参数HMM(GVP-HMM)来显式地对针对变化噪声的最佳、良好匹配的模型参数的连续轨迹进行建模。为了提高传统GVP-HSTRY的泛化能力和计算效率,提出了一种新的GVP-HSTRY模型复杂度控制方法。在局部水平上自动确定高斯均值、方差和模型空间线性变换轨迹的最优多项式次数。分别在Aurora 2和中等词汇量汉语普通话语音识别任务上获得了20%和28%的相对错误率显著降低。一致的性能改进和模型大小压缩60%的相对基线GVP-HMM系统使用统一分配的多项式次数也得到了。
An important part of the acoustic modelling problem for automatic speech recognition (ASR) systems is to handle the mismatch against a target environment created by time-varying external factors such as ambient noise. One possible solution to this problem is to introduce controllability to the underlying acoustic model to allow an instantaneous adaptation to the underlying noise condition. Along this line, the continuous trajectory of optimal, well matched model parameters against the varying noise can be explicitly modelled using, for example, generalized variable parameter HMMs (GVP-HMM). In order to improve the generalization and computational efficiency of conventional GVP-HMMs, this paper investigates a novel model complexity control method for GVP-HMMs. The optimal polynomial degrees of Gaussian mean, variance and model space linear transform trajectories are automatically determined at local level. Significant error rate reductions of 20% and 28% relative were obtained over the multi-style training baseline systems on Aurora 2 and a medium vocabulary Mandarin Chinese speech recognition task respectively. Consistent performance improvements and model size compression of 60% relative were also obtained over the baseline GVP-HMM systems using a uniformly assigned polynomial degree.
DOI: 10.1109/icslp.1996.607807
发表时间: 1996-10
期刊: Proceeding of Fourth International Conference on Spoken Language Processing. ICSLP '96
影响因子: --
作者:
T. Anastasakos;J. McDonough;R. Schwartz;J. Makhoul
通讯作者: T. Anastasakos;J. McDonough;R. Schwartz;J. Makhoul
DOI: 10.1109/icassp.1999.758133
发表时间: 1999-03
期刊: 1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings. ICASSP99 (Cat. No.99CH36258)
影响因子: --
作者:
W. Chou;W. Reichl
通讯作者: W. Chou;W. Reichl
DOI: 10.21437/interspeech.2011-201
发表时间: 2011
期刊: --
影响因子: --
作者:
Ning Cheng;Xunying Liu;Lan Wang
通讯作者: Ning Cheng;Xunying Liu;Lan Wang
DOI: 10.1109/icassp.2003.1198734
发表时间: 2003-04
期刊: 2003 IEEE International Conference on Acoustics, Speech, and Signal Processing, 2003. Proceedings. (ICASSP '03).
影响因子: --
作者:
Xunying Liu;M. Gales;P. Woodland
通讯作者: Xunying Liu;M. Gales;P. Woodland
DOI: 10.1016/j.specom.2007.10.004
发表时间: 2008-04
期刊: Speech Commun.
影响因子: --
作者:
H. Liao;M. Gales
通讯作者: H. Liao;M. Gales