Towards visually-driven speech enhancement for cognitively-inspired multi-modal hearing-aid devices (AV-COGHEAR)
Towards visually-driven speech enhancement for cognitively-inspired multi-modal hearing-aid devices (AV-COGHEAR)
批准号:
EP/M026981/1
负责人:
Amir Hussain
金额:
$53.29万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2015
资助国家:
英国
项目状态:
已结题
起止时间:
2015 至 --
中文摘要
当前的商用助听器使用许多复杂的增强技术来尝试和改善语音信号的质量。然而,今天最好的辅助设备在许多日常情况下都不能很好地发挥作用。特别是,在忙碌的社交场合,有许多相互竞争的语源,它们会失败;如果说话人离听者太远,被噪音淹没,它们就会失败。我们已经找到了一个解决这个问题的机会,方法是制造出可以“看得见”的助听器。这个雄心勃勃的项目旨在开发新一代助听器技术,通过使用摄像头来查看说话者在说什么,从而从噪音中提取语音。该设备的佩戴者将能够将他们的听力集中在目标说话者身上,该设备将过滤掉竞争的声音。这种能力超越了目前的技术,有可能改善数百万听力损失患者的生活质量(仅在英国就超过1000万人)。我们的方法与正常听力是一致的。听众很自然地将耳朵和眼睛的信息结合在一起:我们用眼睛帮助我们听力。在听演讲时,眼睛跟随面部和嘴巴的运动,一个复杂的、多阶段的过程使用这些信息将语音与噪音分开,并填补任何空白。我们的助听器将以大致相同的方式发挥作用。它将利用摄像头的视觉信息(例如,使用类似谷歌眼镜的系统),以及智能地结合音频和视频信息的新颖算法,以提高真实世界噪音环境中的语音质量和可理解性。该项目汇集了一大批研究人员,他们拥有使视听助听器成为可能所需的补充专业知识。该项目将结合斯特林的认知计算小组和谢菲尔德的言语和听力小组开发的视听语音增强的新的对比方法。斯特林方法使用视觉信号来滤除噪声;而谢菲尔德方法使用视觉信号来填补语音中的“空白”。跟踪说话者嘴唇和脸部运动所需的视觉处理将使用斯特林心理学部门开发的革命性条形码表示法。MRC听力研究所(IHR)将提供所需的专业知识,以评估该方法对真实听力损失患者的影响。领先的国际助听器制造商Phonak AG将提供必要的建议和指导,以最大限度地发挥行业影响的潜力。该项目被设计为一系列四个工作包,考虑与该设备设计的每个组件相关的关键研究挑战。谢菲尔德和斯特林的初步工作已经确定了这些问题。其中的挑战包括开发改进的视觉驱动音频分析技术;设计更好的衡量音频和视觉证据权重的指标;开发最佳结合噪音过滤和填补空白的方法的技术。另一个关键挑战是,要使助听器有效,处理过程不能将信号延迟超过10ms。在该项目的最后一年,一个完全集成的软件原型将在一系列现代嘈杂的混响环境中通过听力受损志愿者的听力测试进行临床评估。评估将使用一个新的专门构建的语料库,该语料库将专门为测试这种新的多模式设备而设计。该项目的临床研究合作伙伴,MRC IHR的苏格兰部分,将在整个试验过程中就实验设计和分析方面提供建议。行业领先者Phonak AG将为基准实时听力设备提供建议和技术支持。最终的临床测试原型将提供给整个听力社区作为进一步研究、开发、评估和基准的试验台。
英文摘要
Current commercial hearing aids use a number of sophisticated enhancement techniques to try and improve the quality of speech signals. However, today's best aids fail to work well in many everyday situations. In particular, they fail in busy social situations where there are many competing speech sources; they fail if the speaker is too far from the listener and swamped by noise. We have identified an opportunity to solve this problem by building hearing aids that can 'see'. This ambitious project aims to develop a new generation of hearing aid technology that extracts speech from noise by using a camera to see what the talker is saying. The wearer of the device will be able to focus their hearing on a target talker and the device will filter out competing sound. This ability, which is beyond that of current technology, has the potential to improve the quality of life of the millions suffering from hearing loss (over 10m in the UK alone).Our approach is consistent with normal hearing. Listeners naturally combine information from both their ears and eyes: we use our eyes to help us hear. When listening to speech, eyes follow the movements of the face and mouth and a sophisticated, multi-stage process uses this information to separate speech from the noise and fill in any gaps. Our hearing aid will act in much the same way. It will exploit visual information from a camera (e.g.using a Google Glass like system), and novel algorithms for intelligently combining audio and visual information, in order to improve speech quality and intelligibility in real-world noisy environments. The project is bringing together a critical mass of researchers with the complementary expertise necessary to make the audio-visual hearing-aid possible. The project will combine new contrasting approaches to audio-visual speech enhancement that have been developed by the Cognitive Computing group at Stirling and the Speech and Hearing Group at Sheffield. The Stirling approach uses the visual signal to filter out noise; whereas the Sheffield approach uses the visual signal to fill in 'gaps' in the speech. The vision processing needed to track a speaker's lip and face movement will use a revolutionary 'bar code' representation developed by the Psychology Division at Stirling. The MRC Institute of Hearing Research (IHR) will provide the expertise needed to evaluate the approach on real hearing loss sufferers. Phonak AG, a leading international hearing aid manufacturer, will provide the advice and guidance necessary to maximise potential for industrial impact.The project has been designed as a series of four workpackages that consider the key research challenges related to each component of the device's design. These questions have been identified by preliminary work at Sheffield and Stirling. Among the challenges are developing improved techniques for visually-driven audio-analysis; designing better metrics for weighting audio and visual evidence; developing techniques for optimally combining the noise-filtering and gap-filling approaches. A further key challenge is that, for a hearing aid to be effective, the processing cannot delay the signal by more than 10ms. In the final year of the project a full integrated, software prototype will be clinically evaluated using listening tests with hearing-impaired volunteers in a range of modern noisy reverberant environments. Evaluation will use a new purpose-built speech corpus that will be designed specifically for testing this new class of multimodal device. The project's clinical research partner, the Scottish Section of MRC IHR, will provide advice on the experimental design and analysis aspects throughout the trials. Industry leader Phonak AG will provide advice and technical support for benchmarking real-time hearing devices. The final clinically-tested prototype will be made available to the whole hearing community as a testbed for further research, development, evaluation and benchmarking.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1007/s12559-017-9541-x
发表时间:
2018-01
期刊:
Cognitive Computation
影响因子:
5.4
作者:
[A. Abdullah;A. Hussain;Imtiaz Hussain Khan]
通讯作者:
A. Abdullah;A. Hussain;Imtiaz Hussain Khan
DOI:
10.1109/tetci.2019.2917039
发表时间:
2021-06-01
期刊:
IEEE TRANSACTIONS ON EMERGING TOPICS IN COMPUTATIONAL INTELLIGENCE
影响因子:
5.3
作者:
[Adeel, Ahsan, Gogate, Mandar, Whitmer, William M.]
通讯作者:
Whitmer, William M.
Cognitively Inspired Audiovisual Speech Filtering: Towards an Intelligent, Fuzzy Based, Multimodal, Two-Stage Speech Enhancement System
认知启发的视听语音过滤:走向智能、模糊、多模态、两阶段语音增强系统
DOI:
--
发表时间:
2015
期刊:
影响因子:
--
作者:
[Abel Andrew]
通讯作者:
Abel Andrew
DOI:
10.1007/s00521-022-07839-5
发表时间:
2018-06
期刊:
Neural Computing and Applications
影响因子:
6
作者:
[Wissem Abbes;Zied Kechaou;Amir Hussain;A. Qahtani;Omar Almutiry;Habib Dhahri;A. Alimi]
通讯作者:
Wissem Abbes;Zied Kechaou;Amir Hussain;A. Qahtani;Omar Almutiry;Habib Dhahri;A. Alimi
Context-sensitive neocortical neurons transform the effectiveness and efficiency of neural information processing
上下文敏感的新皮质神经元改变神经信息处理的有效性和效率
DOI:
10.48550/arxiv.2207.07338
发表时间:
2022
期刊:
影响因子:
--
作者:
[Adeel A]
通讯作者:
Adeel A
共 7 条
COG-MHEAR: Towards cognitively-inspired 5G-IoT enabled, multi-modal Hearing Aids
-
批准号:EP/T021063/1
-
项目类别:Research Grant
-
资助金额:$415.26万
-
财政年份:2021
-
负责人:Amir Hussain
-
依托单位:
Dual Process Control Models in the Brain and Machines with Application to Autonomous Vehicle Control
-
批准号:EP/I009310/1
-
项目类别:Research Grant
-
资助金额:$44.93万
-
财政年份:2011
-
负责人:Amir Hussain
-
依托单位:
Industrial CASE Account - Stirling 2009
-
批准号:EP/H501584/1
-
项目类别:Training Grant
-
资助金额:$16.64万
-
财政年份:2009
-
负责人:Amir Hussain
-
依托单位:
Industrial CASE Account - Stirling 2008
-
批准号:EP/G501750/1
-
项目类别:Training Grant
-
资助金额:$8.13万
-
财政年份:2009
-
负责人:Amir Hussain
-
依托单位:
海外基金