LANDMARK-BASED SPEECH RECOGNITION: REPORT OF THE 2004 JOHNS HOPKINS SUMMER WORKSHOP.

LANDMARK-BASED SPEECH RECOGNITION: REPORT OF THE 2004 JOHNS HOPKINS SUMMER WORKSHOP.
复制标题

基于地标的语音识别:2004 年约翰·霍普金斯大学夏季研讨会报告。

DOI:
10.1109/icassp.2005.1415088
复制
发表时间:
2005
期刊:
Proceedings of the ... IEEE International Conference on Acoustics, Speech, and Signal Processing. ICASSP (Conference)
影响因子:
--
通讯作者:
Wang,Tianyu
Wang,Tianyu
中科院分区:
--
文献类型:
--
作者:
Hasegawa-Johnson,Mark;Baker,James;Borys,Sarah;Chen,Ken;Coogan,Emily;Greenberg,Steven;Juneja,Amit;Kirchhoff,Katrin;Livescu,Karen;Mohan,Srividya;Muller,Jennifer;Sonmez,Kemal;Wang,Tianyu

文献摘要

相似文献

描述了三个研究原型语音识别系统,所有这些系统都使用最近开发的人工智能方法(特别是支持向量机(SVM)、动态贝叶斯网络和最大熵分类),以便以 ASR 的形式实现当前人类语音感知和音系学的理论。所有系统都从高维多帧声学到独特特征的转换开始,使用经过训练来检测和分类声学语音标志的 SVM 来实现。然后使用 3 个发音模型之一集成由 SVM 估计的独特特征概率:假设每个单词的规范发音的动态编程算法、发音音系的动态贝叶斯网络实现或使用最大熵分类方法训练的判别发音模型。然后使用对数线性组合将这些模型计算出的对数概率分数与第一遍识别器的网格输出中可用的其他单词分数进行组合,并且所得到的组合分数用于计算第二遍语音识别输出。
Three research prototype speech recognition systems are described, all of which use recently developed methods from artificial intelligence (specifically support vector machines (SVM); dynamic Bayesian networks, and maximum entropy classification) in order to implement, in the form of an ASR, current theories of human speech perception and phonology. All systems begin with a high-D multiframe acoustic-to-distinctive feature transformation, implemented using SVMs trained to detect and classify acoustic phonetic landmarks. Distinctive feature probabilities estimated by the SVMs are then integrated using one of 3 pronunciation models: a dynamic programming algorithm that assumes canonical pronunciation of each word, a dynamic Bayesian network implementation of articulatory phonology, or a discriminative pronunciation model trained using the methods of maximum entropy classification. Log probability scores computed by these models are then combined, using log-linear combination, with other word scores available in the lattice output of a 1st pass recognizer, and the resulting combination score is used to compute a 2nd-pass speech recognition output.