A convolutional neural-network model of human cochlear mechanics and filter tuning for real-time applications

A convolutional neural-network model of human cochlear mechanics and filter tuning for real-time applications
复制标题

DOI:
10.1038/s42256-020-00286-8
复制
发表时间:
2021-02-08
影响因子:
23.8
通讯作者:
Verhulst, Sarah
Verhulst, Sarah
中科院分区:
计算机科学1区
文献类型:
--
作者:
Baby, Deepak;van den Broucke, Arthur;Verhulst, Sarah

文献摘要

被引文献

相似文献

听觉模型通常用作自动语音识别系统的特征提取器,或用作机器人、机器听力和助听器应用的前端。尽管听觉模型可以非常详细地捕获人类听力的生物物理和非线性特性,但这些生物物理模型的计算成本很高,并且无法在实时应用中使用。我们提出了一种混合方法,其中卷积神经网络与计算神经科学相结合,为人类耳蜗力学产生实时端到端模型,包括水平相关的滤波器调整(CoNNear)。 CoNNear 模型在声学语音材料上进行训练,并使用耳蜗力学研究中常用的(看不见的)声音刺激来评估其性能和适用性。 CoNNear 模型准确地模拟了人类耳蜗频率选择性及其对声音强度的依赖性,这是在负语音背景噪声比下实现稳健语音清晰度的基本品质。 CoNNear 架构基于并行和可微分计算,能够实现实时人类表现。这些独特的 CoNNear 功能将使下一代类人机器听力应用成为可能。我们已经开展了大量工作来开发捕获耳朵非线性处理的外围听觉模型。但对于大多数机器听力系统来说,大规模使用所得模型的速度慢得令人望而却步。作者提出了一种卷积神经网络模型,该模型复制了耳蜗信号处理的标志特征,有可能实现实时应用。
Auditory models are commonly used as feature extractors for automatic speech-recognition systems or as front-ends for robotics, machine-hearing and hearing-aid applications. Although auditory models can capture the biophysical and nonlinear properties of human hearing in great detail, these biophysical models are computationally expensive and cannot be used in real-time applications. We present a hybrid approach where convolutional neural networks are combined with computational neuroscience to yield a real-time end-to-end model for human cochlear mechanics, including level-dependent filter tuning (CoNNear). The CoNNear model was trained on acoustic speech material and its performance and applicability were evaluated using (unseen) sound stimuli commonly employed in cochlear mechanics research. The CoNNear model accurately simulates human cochlear frequency selectivity and its dependence on sound intensity, an essential quality for robust speech intelligibility at negative speech-to-background-noise ratios. The CoNNear architecture is based on parallel and differentiable computations and has the power to achieve real-time human performance. These unique CoNNear features will enable the next generation of human-like machine-hearing applications.Extensive work has gone into developing peripheral auditory models that capture the nonlinear processing of the ear. But the resulting models are prohibitively slow to use at scale for most machine hearing systems. The authors present a convolutional neural network model that replicates hallmark features of cochlear signal processing, potentially enabling real-time applications.