The Application of Hidden Markov Models in Speech Recognition

The Application of Hidden Markov Models in Speech Recognition
复制标题

DOI:
10.1561/2000000004
复制
发表时间:
2007-01-01
影响因子:
--
通讯作者:
Young, Steve
Young, Steve
中科院分区:
其他
文献类型:
--
作者:
Gales, Mark;Young, Steve

文献摘要

被引文献

相似文献

隐马尔可夫模型(Hidden Markov Models,简称HMM)为时变谱向量序列建模提供了一个简单有效的框架。因此,目前几乎所有的大词汇量连续语音识别(LVCSR)系统都是基于HMM的,而基于HMM的LVCSR的基本原理是相当简单的,在这些原理的直接实现中涉及的近似和简化假设将导致系统具有较差的准确性和不可接受的敏感性,操作环境的变化。因此,在现代systems的实际应用中的阻碍涉及到相当复杂的。本次审查的目的是首先提出的HMM为基础的LVCSR系统的核心架构,然后描述的各种改进,需要实现国家的最先进的性能。这些改进包括特征投影,改进的协方差建模,判别参数估计,自适应和归一化,噪声补偿和多通道系统组合。最后,本文以广播新闻和会话转录的LVCSR为例,说明了所描述的技术。
Hidden Markov Models (HMMs) provide a simple and effective framework for modelling time-varying spectral vector sequences. As a consequence, almost all present day large vocabulary continuous speech recognition (LVCSR) systems are based on HMMs.Whereas the basic principles underlying HMM-based LVCSR are rather straightforward, the approximations and simplifying assumptions involved in a direct implementation of these principles would result in a system which has poor accuracy and unacceptable sensitivity to changes in operating environment. Thus, the practical application of HMMs in modern systems involves considerable sophistication.The aim of this review is first to present the core architecture of a HMM-based LVCSR system and then describe the various refinements which are needed to achieve state-of-the-art performance. These refinements include feature projection, improved covariance modelling, discriminative parameter estimation, adaptation and normalisation, noise compensation and multi-pass system combination. The review concludes with a case study of LVCSR for Broadcast News and Conversation transcription in order to illustrate the techniques described.