CAREER: Integrating perceptual models of auditory importance into deep learning-based noise-robust speech recognition
CAREER: Integrating perceptual models of auditory importance into deep learning-based noise-robust speech recognition
批准号:
1750383
负责人:
Michael Mandel
金额:
$49.72万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-06-01 至 2023-07-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Hearing is central to human interaction, but the hearing process is not easily observed. The objective of this project is to train models to identify portions of speech utterances that are important to their being correctly identified by human listeners, and to use predictions from these models to make automatic speech recognition (ASR) systems more noise robust by focusing on those regions. The ability to identify important regions of an utterance could significantly advance our understanding of healthy and impaired hearing. Improvements in automatic speech recognition would have broader impacts on the 260 million Americans who use smart phones and the $100 billion ASR industry. The educational portion of this project utilizes examples from speech, language, audio, and music processing to attract and retain students in Brooklyn College's introductory programming course serving a diverse student body along with similar efforts at affiliated high school programs.The team's preliminary results have shown that that some regions of an utterance are more important or useful than others in identifying it by measuring the intelligibility of a given utterance in many different noisy mixtures. This project expands upon these preliminary results in three ways. First it measures ASR auditory importance using the team's existing slow but accurate technique involving random "bubble noise", comparing different ASR variants to each other and to human listeners. Second, it trains a model to predict ASR auditory importance from clean speech using a novel architecture called the bubble cooperative network (BCN) that allows the recognizer to be trained jointly with the BCN to improve performance. Third, it adapts the learned importance predictor to human listeners and uses this human-adapted importance predictor to further refine the ASR models. These tasks should permit the use of utterance-level human responses to directly improve the noise robustness of automatic speech recognition.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/icassp43922.2022.9747003
发表时间:
2021-12
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
[V. Trinh;Hassan Salami Kavaki;Michael Mandel]
通讯作者:
V. Trinh;Hassan Salami Kavaki;Michael Mandel
DOI:
10.21437/interspeech.2020-2883
发表时间:
2020-05
期刊:
影响因子:
--
作者:
[V. Trinh;Michael I. Mandel]
通讯作者:
V. Trinh;Michael I. Mandel
Bubble Cooperative Networks for Identifying Important Speech Cues
用于识别重要语音提示的气泡合作网络
DOI:
10.21437/interspeech.2018-2377
发表时间:
2018
期刊:
Interspeech 2018
影响因子:
--
作者:
[Trinh, Viet Anh, McFee, Brian, Mandel, Michael I]
通讯作者:
Mandel, Michael I
DOI:
10.1109/taslp.2020.3040545
发表时间:
2016-09
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[Michael I. Mandel]
通讯作者:
Michael I. Mandel
The Bubble Noise Technique for Speech Perception Research
用于语音感知研究的气泡噪声技术
DOI:
10.1044/2019_pers-19-00058
发表时间:
2019
期刊:
Perspectives of the ASHA Special Interest Groups
影响因子:
--
作者:
[Mandel, Michael I., Grover, Vikas, Zhao, Mengxuan, Choi, Jiyoung, Shafer, Valerie L.]
通讯作者:
Shafer, Valerie L.
共 7 条
RI: Small: Concatenative Resynthesis for Very High Quality Speech Enhancement
-
批准号:1618061
-
项目类别:Continuing Grant
-
资助金额:$45.0万
-
财政年份:2016
-
负责人:Michael Mandel
-
依托单位:
Is Local Government Representative? a Study of Attitudes
-
批准号:7905335
-
项目类别:Standard Grant
-
资助金额:$2.03万
-
财政年份:1979
-
负责人:Michael Mandel
-
依托单位:
海外基金