课题基金 / 基金详情

RI: Medium: Deep Neural Networks for Robust Speech Recognition through Integrated Acoustic Modeling and Separation

RI: Medium: Deep Neural Networks for Robust Speech Recognition through Integrated Acoustic Modeling and Separation
RI:中:通过集成声学建模和分离实现鲁棒语音识别的深度神经网络
批准号:
1409431
负责人:
Eric Fosler-Lussier
金额:
$79.81万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-06-01 至 2019-05-31

项目摘要

项目成果

Eric Fosler-Lussier的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的十年中,语音识别技术在日常生活中越来越普遍,包括移动个人代理和语音邮件信息转录等应用程序的激增。然而,在存在背景噪声的情况下,这些系统的性能会显著下降;例如,在嘈杂的餐厅或多风的街道上使用语音识别技术可能会很困难,因为语音识别器会将背景噪音与语言内容混淆。噪声补偿通常包括对声信号进行预处理以强调语音信号(即语音分离),然后将处理后的输入输入到识别器中。本项目的创新之处在于将识别和分离系统进行综合训练,使信号的语言内容能够告知分离,反之亦然。考虑到最近深度神经网络(dnn)在语音处理中的复兴的影响,本项目旨在通过整合语音分离和语音识别,探索三个相关领域,使dnn更能抵抗噪声。第一个研究领域旨在通过结合基于dnn的抑制和声学建模,整合跨时间和频率的掩蔽估计,并利用这些信息来改善从噪声输入中重建语音,从而稳定dnn的输入。第二个领域旨在研究更丰富的DNN结构,使用多任务学习技术来指导DNN的构建,以便更好地执行所有任务,并且层具有有意义的结构。最后的研究领域探讨了如何适应给定噪声的DNN声学模型的伪输出。该项目将以整合语音分离和识别为重点,通过测量语音识别性能以及与人类语音感知更密切相关的指标来评估该项目。这将确保本研究产生更广泛的影响,不仅为语音技术提供见解,而且从长远来看有助于下一代听力技术的设计。
英文摘要
Over the last decade, speech recognition technology has become steadily more present in everyday life, as seen by the proliferation of applications including mobile personal agents and transcription of voicemail messages. Performance of these systems, however, degrades significantly in the presence of background noise; for example, using speech recognition technology in a noisy restaurant or on a windy street can be difficult because speech recognizers confuse the background noise with linguistic content. Compensation for noise typically involves preprocessing the acoustic signal to emphasize the speech signal (i.e. speech separation), and then feeding this processed input into the recognizer. The innovative approach in this project is to train the recognition and separation systems in an integrated manner so that the linguistic content of the signal can inform the separation, and vice versa. Given the impact of the recent resurgence of Deep Neural Networks (DNNs) in speech processing, this project seeks to make DNNs more resistant to noise by integrating speech separation and speech recognition, exploring three related areas. The first research area seeks to stabilize input to DNNs by combining DNN-based suppression and acoustic modeling, integrating masking estimates across time and frequency, and using this information to improve reconstruction of speech from noisy input. The second area seeks to examine a richer DNN structure, using multi-task learning techniques to guide the construction of DNNs better at performing all tasks and where layers have meaningful structure. The final research area examines ways to adapt the spurious output of DNN acoustic models given acoustic noise. With the focus of integrating speech separation and recognition, the project will be evaluated both by measuring speech recognition performance, as well as metrics that are more closely related to human speech perception. This will ensure a broader impact of this research by providing insights not only to speech technology but also facilitating the design of next-generation hearing technology in the long run.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Deep Learning Based Complex Spectral Mapping for Multi-Channel Speaker Separation and Speech Enhancement
  • 批准号:
    2125074
  • 项目类别:
    Standard Grant
  • 资助金额:
    $39.06万
  • 财政年份:
    2021
  • 负责人:
    Eric Fosler-Lussier
  • 依托单位:
RI: Small: Early Elementary Reading Verification in Challenging Acoustic Environments
  • 批准号:
    2008043
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2020
  • 负责人:
    Eric Fosler-Lussier
  • 依托单位:
CI-ADDO-NEW: Collaborative Research: The Speech Recognition Virtual Kitchen
  • 批准号:
    1305319
  • 项目类别:
    Standard Grant
  • 资助金额:
    $38.21万
  • 财政年份:
    2013
  • 负责人:
    Eric Fosler-Lussier
  • 依托单位:
CI-P:Collaborative Research:The Speech Recognition Virtual Kitchen
  • 批准号:
    1205424
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.85万
  • 财政年份:
    2012
  • 负责人:
    Eric Fosler-Lussier
  • 依托单位:
海外基金