Automatic Speech Recognition

Automatic Speech Recognition
复制标题

DOI:
10.1007/978-81-322-3972-7_20
复制
发表时间:
2020
期刊:
Fundamentals of Artificial Intelligence
影响因子:
--
通讯作者:
Prof. K. R. Chowdhary
Prof. K. R. Chowdhary
中科院分区:
其他
文献类型:
--
作者:
Prof. K. R. Chowdhary

文献摘要

被引文献

相似文献

自动语音识别(ASR)基本上有两种应用模式:将语音作为口头输入或作为知识源。口语 输入处理像听写系统和导航(事务)系统的应用。使用语音作为知识源具有多媒体索引系统等应用。本章介绍了语音识别过程的各个阶段、ASR的资源、类语音引擎的作用和功能、Julius语音识别引擎、网络语音资源、ASR算法、语言模型和声学模型--类隐马尔可夫模型。介绍了Kaldi语音识别工具包、CMU-Sphinx、HTK、Deep语音识别工具等开源语音识别工具及其使用方法。这些工具具有与高级语言(如C/C++和Python)的接口。接下来是章节总结和练习。
There are basically two application modes for automatic speech recognition (ASR): using speech as spoken input or as knowledge source. Spoken   input addresses applications like dictation systems and navigation (transactional) systems. Using speech as a knowledge source has applications like multimedia indexing systems. The chapter presents the stages of speech recognition process, resources of ASR, role and functions of speech engine—like, Julius speech recognition engine, voice-over web resources, ASR algorithms, language model and acoustic models—like HMM (hidden Markov models). Many open-source tools like—Kaldi speech recognition toolkit, CMU-Sphinx, HTK, and Deep speech tools’ introduction, and guidelines for their usages are presented. These tools have interfaces with high-level languages like C/C++ and Python. The is followed with chapter summary and set of exercises.