Machine Learning Paradigms for Speech Recognition: An Overview

Machine Learning Paradigms for Speech Recognition: An Overview
复制标题

DOI:
10.1109/tasl.2013.2244083
复制
发表时间:
2013-05-01
影响因子:
--
通讯作者:
Li, Xiao
Li, Xiao
中科院分区:
其他
文献类型:
--
作者:
Deng, Li;Li, Xiao

文献摘要

被引文献

相似文献

自动语音识别(ASR)一直是许多机器学习(ML)技术背后的驱动力,包括普遍使用的隐马尔可夫模型、区分学习、结构化序列学习、贝叶斯学习和自适应学习。此外,ML可以而且偶尔确实将ASR作为一个大规模的、现实的应用来严格测试给定技术的有效性,并引发由于语音固有的连续和动态性质而产生的新问题。另一方面,尽管ASR可以在商业上用于某些应用,但这在很大程度上是一个尚未解决的问题--对于几乎所有应用来说,ASR的性能都无法与人类的性能相提并论。现代ML方法论的新见解显示出极大的希望推动ASR技术的最先进水平。这篇综述文章为读者提供了现代ML技术的概述,这些技术在当前使用,并与未来的ASR研究和系统相关。其目的是促进ML和ASR群体之间比过去发生的更多的异花授粉。本文根据已经流行的或可能对ASR技术做出重大贡献的主要ML范例进行组织。本综述介绍和阐述的范例包括:生成性和鉴别性学习;有监督、无监督、半监督和主动学习;适应性和多任务学习;以及贝叶斯学习。这些学习范式是在ASR技术和应用的背景下激发和讨论的。最后,我们介绍和分析了深度学习和稀疏表示学习的最新发展,重点讨论了它们与先进的ASR技术的直接关联。
Automatic Speech Recognition (ASR) has historically been a driving force behind many machine learning (ML) techniques, including the ubiquitously used hidden Markov model, discriminative learning, structured sequence learning, Bayesian learning, and adaptive learning. Moreover, ML can and occasionally does use ASR as a large-scale, realistic application to rigorously test the effectiveness of a given technique, and to inspire new problems arising from the inherently sequential and dynamic nature of speech. On the other hand, even though ASR is available commercially for some applications, it is largely an unsolved problem-for almost all applications, the performance of ASR is not on par with human performance. New insight from modern ML methodology shows great promise to advance the state-of-the-art in ASR technology. This overview article provides readers with an overview of modern ML techniques as utilized in the current and as relevant to future ASR research and systems. The intent is to foster further cross-pollination between the ML and ASR communities than has occurred in the past. The article is organized according to the major ML paradigms that are either popular already or have potential for making significant contributions to ASR technology. The paradigms presented and elaborated in this overview include: generative and discriminative learning; supervised, unsupervised, semi-supervised, and active learning; adaptive and multi-task learning; and Bayesian learning. These learning paradigms are motivated and discussed in the context of ASR technology and applications. We finally present and analyze recent developments of deep learning and learning with sparse representations, focusing on their direct relevance to advancing ASR technology.