CAREER: Optimizing Human Speech Perception in Noisy Environments with User-Guided Machine Learning
CAREER: Optimizing Human Speech Perception in Noisy Environments with User-Guided Machine Learning
批准号:
1942718
负责人:
Donald Williamson
金额:
$55.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-06-01 至 2022-09-30
中文摘要
在每年近200亿次的视频电话会议期间,以及数百万助听器用户,不想要的背景噪音往往会阻碍通过设备进行的交流。人们开发了一些方法来消除不需要的噪声,但不幸的是,它们在许多实际环境中并不能很好地执行。随后,去噪方法往往提供低质量和难以理解的收听体验,这导致用户不满意和沮丧。该学院早期承运人开发项目将开发降噪和评估方法,以解决这些问题,从而改善用户的听力体验。经常使用数字手段(如语音会议和助听器)进行人与人之间交流的个人和公司将是这项工作的主要受益者。这项研究产生的数据和算法将被提供给来自不同和跨学科领域的科学家和研究人员。此外,基于这项研究的教育活动将被整合到各种努力中,以增加这些研究领域中代表性不足的参与者的数量。该项目的主要目标是开发用户引导的机器学习算法,从而在真实世界嘈杂的环境中改善听力体验。在包含许多相互竞争的说话者的环境中,降噪系统会无意中删除或保留非预期的语音信号。拟议的研究活动将通过(1)开发识别特定用户想要听到的语音信号的多模式计算方法来解决这一问题。研究人员通常使用计算评估指标来评估绩效,但它们并不总是与个人用户情绪相关,这意味着调查人员的评估结果不准确。该项目将(2)开发一个有效的界面,用于捕捉和预测用户对质量和可理解性的短期评估。模拟语音数据和真实世界的语音数据在说话人、噪声和环境特征方面存在差异,但现有的降噪方法不能动态适应这些差异。这是一个主要缺点,因为部署的降噪系统将遇到未知的扬声器和噪音。研究人员将(3)开发一类新的用户引导的机器学习算法,该算法利用真实和预测的用户评估近乎实时地进行系统优化。成功完成这些任务将有助于更好地理解语音感知并提高降噪系统的可用性。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Unwanted background noise often hinders device-mediated communication during the nearly 20 billion yearly video conference calls and for millions of hearing aid users. Approaches are developed to remove unwanted noise, but unfortunately, they do not perform well in many real environments. Subsequently, the noise-removal approaches often provide low quality and unintelligible listening experiences, which results in dissatisfied and frustrated users. This Faculty Early Carrer Development project will develop noise-reduction and assessment approaches that address these issues, resulting in improved listening experiences for users. Individuals and companies that regularly use digital means (e.g. voice conferencing and hearing aids) for person-to-person communication will be major beneficiaries of this work. The data and algorithms that result from this research will be made available to benefit scientists and researchers from diverse and interdisciplinary fields. Additionally, educational activities based on this research will be integrated into various efforts to increase the number of underrepresented participants in these research areas.The main objective of this project is to develop user-guided machine-learning algorithms that result in improved listening experiences in real-world noisy environments. In environments that contain many competing talkers, noise-reduction systems inadvertently remove or retain unintended speech signals. The proposed research activities will address this by (1) developing multi-modal computational approaches that identify the speech signal that a specific user wants to hear. Computational assessment metrics are generally used by researchers to assess performance, but they do not always correlate with individual user sentiment, meaning investigators have inaccurate assessment results. This project will (2) develop an effective interface for capturing and predicting short-time user assessment of quality and intelligibility. Simulated and real-world speech data differ in terms of speaker, noise and environmental characteristics, but current noise-reduction approaches are incapable of adapting to these differences on the fly. This is a major shortcoming as deployed noise-reduction systems will encounter unknown speakers and noises. The investigator will (3) develop a novel class of user-guided machine learning algorithms that utilize true and predicted user assessment in near-real time for system optimization. Successfully completing these tasks will help better understand speech perception and increase the usability of noise-reduction systems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1109/taslp.2023.3328282
发表时间:
2023-03
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
作者:
[Khandokar Md. Nayem;D. Williamson]
通讯作者:
Khandokar Md. Nayem;D. Williamson
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
从感知语音质量角度:频段对噪声的鲁棒性
DOI:
10.1121/10.0025272
发表时间:
2024
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
作者:
[Fan, Junyi, Williamson, Donald S.]
通讯作者:
Williamson, Donald S.
CAREER: Optimizing Human Speech Perception in Noisy Environments with User-Guided Machine Learning
-
批准号:2235228
-
项目类别:Continuing Grant
-
资助金额:$55.0万
-
财政年份:2022
-
负责人:Donald Williamson
-
依托单位:
CRII: RI: Towards Human-Level Assessment of Speech Quality and Intelligibility in Real-World Environments
-
批准号:1755844
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2018
-
负责人:Donald Williamson
-
依托单位:
海外基金