课题基金 / 基金详情

项目摘要

项目成果

JAMES M COUGHLAN的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):基于视频的听力和视力丧失者语音增强项目概述预计到2030年,美国65岁以上人口将占总人口的20%以上。听力和视力的丧失自然伴随着衰老的过程。听力损失的人可以从观察说话人的视觉线索中获益,比如嘴唇的形状和面部表情,从而大大提高他们理解言语的能力。然而,视力丧失的人不能利用这些视觉线索,并且很难理解言语,特别是在嘈杂的环境中。此外,视力正常的人可以使用视觉信息来识别群体中的说话者,这使他们能够专注于这个人。这对正在使用扩音器或助听器等设备的听力损失患者大有裨益。然而,需要向视力丧失的用户提供这些扬声器信息,以便最佳地使用这些设备。我们建议开发一种原型装置,用于清除目标说话人的语音信号,提高听力和视力丧失者在日常情况下的语音理解能力。为了完成这项任务,我们需要利用视觉线索,迄今为止,在为听力损失人士设计辅助技术时,这些线索在很大程度上被忽视了。我们的第一个目标是学习与目标语音信号相关的与说话人无关的视觉线索,并使用这些视听线索来设计语音增强算法,这些算法在嘈杂的日常环境中比目前仅利用音频信号的方法表现得更好。我们将利用摄像机和计算机视觉方法设计先进的数字信号处理技术,以增强通过麦克风录制的目标语音信号。我们的第二个目标是利用视频和音频信号来检测和有效地定位可见说话人。然后,有关感兴趣的说话者位置的信息可用于有效地执行说话者分离,并提供给用户。最后,我们的目标是在便携式原型系统上实现这些开发的算法。我们将通过实际情况和实验室条件下的用户实验来测试该系统的性能并改进用户界面。最终产品将显示将多种模式纳入感官辅助设备的可行性和重要性,并为未来的研究和开发工作奠定基础。
英文摘要
DESCRIPTION (provided by applicant): Video-based Speech Enhancement for Persons with Hearing and Vision Loss Project Summary It is estimated that by 2030, the number of people in the United States over the age of 65 will account for over 20% of the total population. Hearing and vision loss naturally accompanies the aging process. Persons with hearing loss can benefit from observing the visual cues from a speaker such as the shape of the lips and facial expression to greatly improve their ability to comprehend speech. However, persons with vision loss cannot make use of these visual cues, and have a harder time understanding speech, especially in noisy environments. Furthermore, people with normal vision can use visual information to identify a speaker in a group, which allows them to focus on this person. This can greatly benefit a person with hearing loss who may be using a device such as a sound amplifier or a hearing aid. A user with vision loss, however, needs to be provided with this speaker information to make optimal use of such devices. We propose developing a prototype device that will clean the speech signal from a target speaker and improve speech comprehension for persons with hearing and vision loss in everyday situations. In order to accomplish this task, we need to harness the visual cues that have so far largely been ignored in the design of assistive technolo- gies for persons with hearing loss. Our first aim is to learn speaker-independent visual cues that are associated with the target speech signal, and use these audio-visual cues to design speech enhancement algorithms that perform much better in noisy everyday environment than current methods which only utilize the audio signal. We will utilize a video camera and computer vision methods to design advanced digital signal processing techniques to enhance the target speech signals recorded through a microphone. Our second aim is to use the video and audio signals to detect and efficiently localize the visible speaker. The information regarding the location of the speaker of interest can then be used to efficiently perform speaker separation, as well as be provided to the user. Finally, we aim to implement these developed algorithms on a portable prototype system. We will test the performance of this system and improve the user-interface through user experiments in real-world situations as well as laboratory conditions. The end product will show the feasibility and importance of incorporating multiple modalities into sensory assistive devices, and set the stage for future research and development efforts.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
Leveraging Maps and Computer Vision to Support Indoor Navigation for Blind Travelers
Leveraging Maps and Computer Vision to Support Indoor Navigation for Blind Travelers
Leveraging Maps and Computer Vision to Support Indoor Navigation for Blind Travelers
Enabling Audio-Haptic Interaction with Physical Objects for the Visually Impaired
海外基金