LANDMARK-BASED SPEECH RECOGNITION: REPORT OF THE 2004 JOHNS HOPKINS SUMMER WORKSHOP.
LANDMARK-BASED SPEECH RECOGNITION: REPORT OF THE 2004 JOHNS HOPKINS SUMMER WORKSHOP.
复制标题
基于地标的语音识别:2004 年约翰·霍普金斯大学夏季研讨会报告。
DOI:
10.1109/icassp.2005.1415088
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
Wang,Tianyu
中科院分区:
文献类型:
--
作者:
Hasegawa-Johnson,Mark;Baker,James;Borys,Sarah;Chen,Ken;Coogan,Emily;Greenberg,Steven;Juneja,Amit;Kirchhoff,Katrin;Livescu,Karen;Mohan,Srividya;Muller,Jennifer;Sonmez,Kemal;Wang,Tianyu
Three research prototype speech recognition systems are described, all of which use recently developed methods from artificial intelligence (specifically support vector machines (SVM); dynamic Bayesian networks, and maximum entropy classification) in order to implement, in the form of an ASR, current theories of human speech perception and phonology. All systems begin with a high-D multiframe acoustic-to-distinctive feature transformation, implemented using SVMs trained to detect and classify acoustic phonetic landmarks. Distinctive feature probabilities estimated by the SVMs are then integrated using one of 3 pronunciation models: a dynamic programming algorithm that assumes canonical pronunciation of each word, a dynamic Bayesian network implementation of articulatory phonology, or a discriminative pronunciation model trained using the methods of maximum entropy classification. Log probability scores computed by these models are then combined, using log-linear combination, with other word scores available in the lattice output of a 1st pass recognizer, and the resulting combination score is used to compute a 2nd-pass speech recognition output.