课题基金 / 基金详情

RI: Medium: Collaborative Research: Multilingual Gestural Models for Robust Language-Independent Speech Recognition

RI: Medium: Collaborative Research: Multilingual Gestural Models for Robust Language-Independent Speech Recognition
RI:媒介:协作研究:用于鲁棒语言无关语音识别的多语言手势模型
批准号:
1162525
负责人:
Carol Espy-Wilson
金额:
$23.49万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-10-01 至 2016-09-30

项目摘要

项目成果

Carol Espy-Wilson的其他基金

相似基金

相关文献

中文摘要
翻译
当前最先进的自动语音识别(ASR)系统通常将语音建模为一串声学定义的音素,并使用诸如三音素或五音素之类的语境化音素单元来对由于协同发音而产生的语境影响进行建模。这样的声学模型可能会受到数据稀疏性的影响,并且可能无法适当地捕捉协同发音,因为三音素或五音素的上下文影响的范围不灵活。然而,在小词汇量的情况下,研究表明,根据声学估计发音手势并将这些手势纳入ASR过程的ASR系统可以更好地模拟协同发音,并且对噪声更健壮。当前的项目调查了估计的发音手势在大词汇量自动语音识别中的使用。语音信号的手势表示最初是使用语音产生的任务动态模型从声学波形创建的。这些数据然后被用来训练发音手势识别的自动模型,其中发音手势在基于手势的ASR系统中充当子词单位。这项工作的主要目标是评估使用美国英语(AE)的基于大词汇量手势的ASR系统的性能。这项研究的广泛影响有三个方面:(1)创建了一个包含声学波形及其发音表示的大词汇量美国英语(AE)语音数据库;(2)引入了新的机器学习技术来模拟声学波形的发音表示;(3)开发了一个以发音表示为子词单位的大词汇量自动语音识别系统。由拟议的项目产生的用于声发射的稳健和准确的ASR系统将有效地处理语音变化,从而显著增强声发射中人与机器之间的沟通和协作,并承诺将该方法推广到多种语言。所获得的知识和开发的系统将有助于发音特征在语音处理中的广泛应用,并将有可能改变自动语音识别、语音介导人机交互和语言之间的自动翻译领域。跨学科合作将促进参与的教师、研究人员、研究生和本科生的跨学科学习环境,因此,这种合作将导致语音建模和算法开发方面的强化培训产生更广泛的影响。最后,拟议的工作将产生一套数据库和工具,这些数据库和工具将被传播,为广大研究和教育界服务。
英文摘要
Current state-of-the-art automatic speech recognition (ASR) systems typically model speech as a string of acoustically-defined phones and use contextualized phone units, such as tri-phones or quin-phones to model contextual influences due to coarticulation. Such acoustic models may suffer from data sparsity and may fail to capture coarticulation appropriately because the span of a tri- or quin-phone's contextual influence is not flexible. In a small vocabulary context, however, research has shown that ASR systems which estimate articulatory gestures from the acoustics and incorporate these gestures in the ASR process can better model coarticulation and are more robust to noise. The current project investigates the use of estimated articulatory gestures in large vocabulary automatic speech recognition. Gestural representations of the speech signal are initially created from the acoustic waveform using the Task Dynamic model of speech production. These data are then used to train automatic models for articulatory gesture recognition where the articulatory gestures serve as subword units in the gesture-based ASR system. The main goal of the proposed work is to evaluate the performance of a large-vocabulary gesture-based ASR system using American English (AE). The gesture-based system will be compared to a set of competitive state-of-the-art recognition systems in term of word and phone recognition accuracies, both under clean and noisy acoustic background conditions.The broad impact of this research is threefold: (1) the creation of a large vocabulary American English (AE) speech database containing acoustic waveforms and their articulatory representations, (2) the introduction of novel machine learning techniques to model articulatory representations from acoustic waveforms, and (3) the development of a large vocabulary ASR system that uses articulatory representation as subword units. The robust and accurate ASR system for AE resulting from the proposed project will deal effectively with speech variability, thereby significantly enhancing communication and collaboration between people and machines in AE, and with the promise to generalize the method to multiple languages. The knowledge gained and the systems developed will contribute to the broad application of articulatory features in speech processing, and will have the potential to transform the fields of ASR, speech-mediated person-machine interaction, and automatic translation among languages. The interdisciplinary collaboration will facilitate a cross-disciplinary learning environment for the participating faculty, researchers, graduate students and undergraduate students Thus, this collaboration will result in the broader impact of enhanced training in speech modeling and algorithm development. Finally, the proposed work will result in a set of databases and tools that will be disseminated to serve the research and education community at large.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Estimating Articulatory Constriction Place and Timing from Speech Acoustics
  • 批准号:
    2141413
  • 项目类别:
    Standard Grant
  • 资助金额:
    $24.54万
  • 财政年份:
    2022
  • 负责人:
    Carol Espy-Wilson
  • 依托单位:
SCH: INT: Collaborative Research: Using Multi-Stage Learning to Prioritize Mental Health
  • 批准号:
    2124270
  • 项目类别:
    Standard Grant
  • 资助金额:
    $84.24万
  • 财政年份:
    2021
  • 负责人:
    Carol Espy-Wilson
  • 依托单位:
Speech for Robotics
  • 批准号:
    1941541
  • 项目类别:
    Standard Grant
  • 资助金额:
    $5.0万
  • 财政年份:
    2019
  • 负责人:
    Carol Espy-Wilson
  • 依托单位:
Collaborative Research: Effects of production variability on the acoustic consequences of coordinated articulatory gestures
  • 批准号:
    1436600
  • 项目类别:
    Standard Grant
  • 资助金额:
    $13.24万
  • 财政年份:
    2014
  • 负责人:
    Carol Espy-Wilson
  • 依托单位:
海外基金