课题基金 / 基金详情

EAGER: How does deep learning improve speech recognition accuracy?

EAGER: How does deep learning improve speech recognition accuracy?
EAGER:深度学习如何提高语音识别准确性?
批准号:
1450916
负责人:
Steven Wegmann
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2016-08-31

项目摘要

项目成果

Steven Wegmann的其他基金

相似基金

相关文献

中文摘要
翻译
普遍和准确的自动语音识别有可能以许多积极的方式改变社会,其中最重要的是为那些难以甚至无法使用键盘与计算机交互的人提供更好的信息获取途径,例如老年人、身体残疾人士或视力受损人士。每天都有数百万人使用基于该技术的应用程序来解决通过语音与机器交互最自然完成的问题。然而,这些应用中最成功的应用范围总是相当有限,因为语音识别虽然有用,但可能令人沮丧地不可靠。例如,尽管在拥挤的房间里有很大的背景噪音,电话频道严重失真,或者他们共同语言中的口音差异很大,但即使是这些问题中更温和的例子也会使语音识别系统完全脱轨。这个早期探索性研究基金(EAGER)支持一个项目,该项目的短期目标是深入地、定量地理解为什么几乎所有语音识别器中使用的方法都是如此脆弱。长期目标是通过开发不那么脆弱的方法来利用这种理解,从而使更准确的语音识别具有更广泛的适用性。这项探索性研究将首先发现为什么多层感知器(MLP)有时可以提高语音识别的准确性,其次,使用这些诊断见解来选择更好的MLP架构,第三,发布软件,以便其他人可以利用开发的方法。mlp的使用在过去十年中出现了显著的复苏,特别是最近开发的“深度”架构。在语音识别领域,mlp的两个应用都显著提高了大词汇量语音识别的准确率。这些应用程序中的每一个都在标准语音识别机器中工作,该机器使用隐马尔可夫模型(hmm)来模拟声学,模型输入(特征)的梅尔频率倒谱系数(MFCCs)和隐藏状态边缘分布的多元正态。第一个应用程序通过使用MLP从数据中学习到的新特征来增加标准模型输入,从而对标准机器进行了相对较小的调整。第二个应用程序对标准机制进行了更大的改变,用一个对边缘状态后验建模的MLP替换了隐藏状态的边缘分布集合。探索性研究将发现基于mlp的特征能够有效提高基于hmm的语音识别准确率的基本机制。本研究建立在ICSI实验室之前的工作基础上,该实验室使用模拟和新颖的采样过程来量化主要HMM假设对语音识别准确性的影响。最近关于语音识别的mlp的其他研究要么集中在实现上,即如何实际提高语音识别的准确性,要么集中在理论的渐近结果上。虽然这项研究显然很重要,但它主要是通过试验和错误进行的,特别是,它没有解决围绕这些mlp应用如何实际提高语音识别准确性的有趣科学问题。在短期内,对后一个问题的更深入理解应该会导致语音识别准确性的进一步提高,从长远来看,能够开发出比HMM更合适和成功的语音识别模型,这将是该领域的革命性进步。
英文摘要
Pervasive and accurate automatic speech recognition has the potential to transform society in many positive ways, not the least of which is providing better access to information for those who find it difficult or even impossible to interact with computers using a keyboard: e.g. the elderly, the physically disabled, or the vision impaired. Every day millions of people use applications based on this technology to solve problems that are most naturally accomplished by interacting with machines via voice. However, the most successful of these applications have always been rather limited in scope, because, although useful, speech recognition can be frustratingly unreliable. For example, human beings are easily able to understand one another despite loud background noise in a crowded room, severe distortion over a telephone channel, or wide variation in accents within their common language, but even much milder examples of these problems will completely derail a speech recognition system. This EArly Grant for Exploratory Research (EAGER) supports a project whose short term goal is to understand in a deep, quantitative way why methodology used in nearly all speech recognizers is so brittle. The long term goal is to leverage this understanding by developing less brittle methodology that will enable more accurate speech recognition with a wider scope of applicability. This exploratory study will, first, discover why multilayer perceptrons (MLPs) can sometimes improve speech recognition accuracy, second, use these diagnostic insights to select better MLP architectures, and, third, release software so that others can leverage the developed methods. The use of MLPs has staged a remarkable resurgence in the last decade, in particular the "deep" architectures developed recently. In the field of speech recognition, there are two applications of MLPs that have significantly improved large vocabulary speech recognition accuracy. Each of these applications work within the standard speech recognition machinery, which uses hidden Markov models (HMMs) to model the acoustics, mel-frequency cepstral coefficients (MFCCs) for the models' inputs (features), and multivariate normals for the hidden states' marginal distributions. The first application makes a relatively minor adjustment to the standard machinery by augmenting the standard model inputs with new features learned from data using a MLP. The second application makes a more substantial change to the standard machinery by replacing the collection of hidden states' marginal distributions with a single MLP that models the marginal state posteriors. The research in the exploratory study will discover the basic mechanisms that the MLP-based features use to substantially improve HMM-based speech recognition accuracy. This research builds upon the previous work in the ICSI lab that used simulation and a novel sampling process to quantify the impacts that the major HMM assumptions have on speech recognition accuracy.Other recent research on MLPs for speech recognition has either concentrated on implementation, i.e., how to actually improve speech recognition accuracy, or on theoretical asymptotic results. While this research is obviously important, it has proceeded largely by trial and error and, in particular, it has not addressed the interesting scientific questions surrounding how these applications of MLPs actually improve speech recognition accuracy. A deeper understanding of this latter question should, in the short term, lead to further improvements in speech recognition accuracy and, in the long term, enable the development of more suitable and successful models for speech recognition than the HMM, which would be a transformative advance in the field.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RI: Small: Exploratory Data Analysis for Speech Recognition
海外基金