COG-MHEAR: Towards cognitively-inspired 5G-IoT enabled, multi-modal Hearing Aids
COG-MHEAR: Towards cognitively-inspired 5G-IoT enabled, multi-modal Hearing Aids
批准号:
EP/T021063/1
负责人:
Amir Hussain
金额:
$415.26万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
目前,只有40%的人可以从助听器(HA)中受益,并且大多数拥有HA设备的人并不经常使用它们。使用可见的HAs(“害怕看起来老”)有社会耻辱感,它们需要大量有意识的努力来专注于不同的声音和扬声器,并且只有有限的使用语音增强-使口语(这通常是人们听力的最重要方面)更容易区分。仅仅把一切都弄得更大声是不够的!为了在2050年之前改变听力保健,我们的目标是完全重新思考HA的设计方式。我们的变革性方法-第一次-借鉴了正常听力的认知原则。听者自然地将来自耳朵和眼睛的信息联合收割机结合起来:我们用眼睛来帮助我们听。我们将创造“多模态”辅助设备,不仅可以放大声音,还可以根据上下文同时使用从一系列传感器收集的信息来提高语音清晰度。例如,关于一个人所说的话的大量信息以视觉信息、说话者的嘴唇的运动、手势等来传达。这被当前的商业HA所忽略,并且可以被馈送到语音增强过程中。我们还可以使用可穿戴传感器(嵌入HA本身)来估计听力努力及其对人的影响,并使用它来判断语音增强过程是否真的有帮助。创建这些多模态“视听”HA提出了许多需要整体解决的艰巨技术挑战。传统上,利用嘴唇的动作需要摄像机拍摄演讲者,这就带来了隐私问题。我们可以通过在收集到数据后立即对其进行加密来克服其中的一些问题,并且我们将开创新的方法来处理和理解视频数据,同时保持加密状态。我们的目标是永远不访问原始视频数据,但仍然将其用作有用的信息来源。为了补充这一点,我们还将研究不使用视频馈送的远程唇阅读方法,而是探索使用无线电信号进行远程监控。添加这些新的传感器和所需的处理来理解所产生的数据将对HA设备造成显著的额外功率和小型化负担。我们将需要使我们复杂的视觉和声音处理算法以最小的功耗和最小的延迟运行,并将通过专门的硬件实现来实现这一目标,加速关键的处理步骤。长远来说,我们的目标是所有处理工作均由医管局自行进行,而数据只会存放在有关人士的本地,以保障其隐私。在短期内,一些处理将需要在云中完成(因为它太耗电了),我们将创建新的非常低延迟(<10 ms)的云基础设施接口,以避免在“看到”一个单词和听到它之间的延迟。我们还计划利用柔性电子产品(电子皮肤)和天线设计的进步,使整个单元尽可能小,谨慎和可用。与HA制造商,临床医生和最终用户的预先设计和共同生产将是上述所有内容的核心,指导在设计,优先级和形状因素方面做出的所有决策。我们强大的用户群,其中包括Sonova,诺基亚/贝尔实验室,聋人苏格兰和听力损失行动将有助于最大限度地发挥我们雄心勃勃的研究计划的影响。我们的工作成果将完全集成,软件和硬件原型,将在一系列现代嘈杂的混响环境中使用听力和可懂度测试对听力受损的志愿者进行临床评估。我们雄心勃勃的愿景是否成功,将取决于我们的示范计划所提出的根本性进步将如何重塑未来十年的医管局格局。
英文摘要
Currently, only 40% of people who could benefit from Hearing Aids (HAs) have them, and most people who have HA devices don't use them often enough. There is social stigma around using visible HAs ('fear of looking old'), they require a lot of conscious effort to concentrate on different sounds and speakers, and only limited use is made of speech enhancement - making the spoken words (which are often the most important aspect of hearing to people) easier to distinguish. It is not enough just to make everything louder!To transform hearing care by 2050, we aim to completely re-think the way HAs are designed. Our transformative approach - for the first time - draws on the cognitive principles of normal hearing. Listeners naturally combine information from both their ears and eyes: we use our eyes to help us hear. We will create "multi-modal" aids which not only amplify sounds but contextually use simultaneously collected information from a range of sensors to improve speech intelligibility. For example, a large amount of information about the words said by a person is conveyed in visual information, in the movements of the speaker's lips, hand gestures, and similar. This is ignored by current commercial HAs and could be fed into the speech enhancement process. We can also use wearable sensors (embedded within the HA itself) to estimate listening effort and its impact on the person, and use this to tell whether the speech enhancement process is actually helping or not.Creating these multi-modal "audio-visual" HAs raises many formidable technical challenges which need to be tackled holistically. Making use of lip movements traditionally requires a video camera filming the speaker, which introduces privacy questions. We can overcome some of these questions by encrypting the data as soon as it is collected, and we will pioneer new approaches for processing and understanding the video data while it stays encrypted. We aim to never access the raw video data, but still to use it as a useful source of information. To complement this, we will also investigate methods for remote lip reading without using a video feed, instead exploring the use of radio signals for remote monitoring. Adding in these new sensors and the processing that is required to make sense of the data produced will place a significant additional power and miniaturization burden on the HA device. We will need to make our sophisticated visual and sound processing algorithms operate with minimum power and minimum delay, and will achieve this by making dedicated hardware implementations, accelerating the key processing steps. In the long term, we aim for all processing to be done in the HA itself - keeping data local to the person for privacy. In the shorter term, some processing will need to be done in the cloud (as it is too power intensive) and we will create new very low latency (<10ms) interfaces to cloud infrastructure to avoid delays between when a word is "seen" being spoken and when it is heard. We also plan to utilize advances in flexible electronics (e-skin) and antenna design to make the overall unit as small, discreet and usable as possible. Participatory design and co-production with HA manufacturers, clinicians and end-users will be central to all of the above, guiding all of the decisions made in terms of design, prioritisation and form factor. Our strong User Group, which includes Sonova, Nokia/Bell Labs, Deaf Scotland and Action on Hearing Loss will serve to maximise the impact of our ambitious research programme. The outcomes of our work will be fully integrated, software and hardware prototypes, that will be clinically evaluated using listening and intelligibility tests with hearing-impaired volunteers in a range of modern noisy reverberant environments. The success of our ambitious vision will be measured in terms of how the fundamental advancements posited by our demonstrator programme will reshape the HA landscape over the next decade.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1007/s00521-022-07839-5
发表时间:
2018-06
期刊:
Neural Computing and Applications
影响因子:
6
作者:
[Wissem Abbes;Zied Kechaou;Amir Hussain;A. Qahtani;Omar Almutiry;Habib Dhahri;A. Alimi]
通讯作者:
Wissem Abbes;Zied Kechaou;Amir Hussain;A. Qahtani;Omar Almutiry;Habib Dhahri;A. Alimi
Cooperation Is All You Need
合作就是您所需要的
DOI:
10.48550/arxiv.2305.10449
发表时间:
2023
期刊:
影响因子:
--
作者:
[Adeel A]
通讯作者:
Adeel A
Privacy-Preserving British Sign Language Recognition Using Deep Learning
使用深度学习保护隐私的英国手语识别
DOI:
10.36227/techrxiv.19170257.v1
发表时间:
2022
期刊:
影响因子:
--
作者:
[Abbasi Q]
通讯作者:
Abbasi Q
DOI:
10.1016/j.inffus.2020.10.003
发表时间:
2021-03-01
期刊:
INFORMATION FUSION
影响因子:
18.6
作者:
[Al-Ghadir, Abdulrahman I., Azmi, Aqil M., Hussain, Amir]
通讯作者:
Hussain, Amir
Live Demonstration: Unlocking the Potential of Two-Point Neuronal Cells for Energy-Efficient Training of Deep Networks
现场演示:释放两点神经元细胞的潜力,实现深度网络的节能训练
DOI:
10.1109/iscas46773.2023.10181836
发表时间:
2023
期刊:
影响因子:
--
作者:
[Adeel A]
通讯作者:
Adeel A
共 7 条
Towards visually-driven speech enhancement for cognitively-inspired multi-modal hearing-aid devices (AV-COGHEAR)
-
批准号:EP/M026981/1
-
项目类别:Research Grant
-
资助金额:$53.29万
-
财政年份:2015
-
负责人:Amir Hussain
-
依托单位:
Dual Process Control Models in the Brain and Machines with Application to Autonomous Vehicle Control
-
批准号:EP/I009310/1
-
项目类别:Research Grant
-
资助金额:$44.93万
-
财政年份:2011
-
负责人:Amir Hussain
-
依托单位:
Industrial CASE Account - Stirling 2009
-
批准号:EP/H501584/1
-
项目类别:Training Grant
-
资助金额:$16.64万
-
财政年份:2009
-
负责人:Amir Hussain
-
依托单位:
Industrial CASE Account - Stirling 2008
-
批准号:EP/G501750/1
-
项目类别:Training Grant
-
资助金额:$8.13万
-
财政年份:2009
-
负责人:Amir Hussain
-
依托单位: