Automatic Speech Recognition
Automatic Speech Recognition
复制标题
DOI:
10.1007/978-81-322-3972-7_20
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Prof. K. R. Chowdhary
中科院分区:
文献类型:
--
作者:
Prof. K. R. Chowdhary
There are basically two application modes for automatic speech recognition (ASR): using speech as spoken input or as knowledge source. Spoken input addresses applications like dictation systems and navigation (transactional) systems. Using speech as a knowledge source has applications like multimedia indexing systems. The chapter presents the stages of speech recognition process, resources of ASR, role and functions of speech engine—like, Julius speech recognition engine, voice-over web resources, ASR algorithms, language model and acoustic models—like HMM (hidden Markov models). Many open-source tools like—Kaldi speech recognition toolkit, CMU-Sphinx, HTK, and Deep speech tools’ introduction, and guidelines for their usages are presented. These tools have interfaces with high-level languages like C/C++ and Python. The is followed with chapter summary and set of exercises.