Graphical model architectures for speech recognition

Graphical model architectures for speech recognition
复制标题

DOI:
10.1109/msp.2005.1511827
复制
发表时间:
2005-09-01
影响因子:
14.9
通讯作者:
Bartels, C
Bartels, C
中科院分区:
工程技术1区
文献类型:
--
作者:
Bilmes, JA;Bartels, C

文献摘要

被引文献

相似文献

本文讨论了在语音识别中使用图形模型的基础,如J. R。Deller et al.(1993),X. D. Huang等人(2001),F. Jelinek(19970,L. R. Rabiner和B。-H. Juang(1993)和S. Young等人(1990)详细介绍了一些较为成功的案例。我们的讨论使用动态贝叶斯网络(DBN)和使用图形模型工具包(GMTK)基本模板的DBN扩展,这是一种更适合语音和语言系统的动态图形模型表示。虽然本文集中讨论语音识别,但应该注意的是,这里提出的许多想法也适用于自然语言处理和一般的时间序列分析。
This article discusses the foundations of the use of graphical models for speech recognition as presented in J. R. Deller et al. (1993), X. D. Huang et al. (2001), F. Jelinek (19970, L. R. Rabiner and B. -H. Juang (1993) and S. Young et al. (1990) giving detailed accounts of some of the more successful cases. Our discussion employs dynamic Bayesian networks (DBNs) and a DBN extension using the Graphical Model Toolkit's (GMTK's) basic template, a dynamic graphical model representation that is more suitable for speech and language systems. While this article concentrates on speech recognition, it should be noted that many of the ideas presented here are also applicable to natural language processing and general time-series analysis.