Adaptive temporal encoding leads to a background-insensitive cortical representation of speech.

Adaptive temporal encoding leads to a background-insensitive cortical representation of speech.
复制标题

DOI:
10.1523/jneurosci.5297-12.2013
复制
发表时间:
2013-03-27
期刊:
The Journal of neuroscience : the official journal of the Society for Neuroscience
影响因子:
--
通讯作者:
Simon JZ
Simon JZ
中科院分区:
其他
文献类型:
--
作者:
Ding N;Simon JZ

文献摘要

被引文献

相似文献

语音识别是非常强大的听力背景,即使当背景声音的能量与语音的强烈重叠。然而,大脑如何将受损的声学信号转换为适用于语音识别的可靠神经表示仍然是难以捉摸的。在这里,我们假设这种转换是在听觉皮层的水平,通过自适应神经编码,我们测试的假设记录,使用脑磁图(MEG),人类受试者听一个叙述的故事的神经反应。与语音具有最大声学重叠的频谱匹配平稳噪声以不同的强度级别混合。尽管这种噪声引起的严重的声学干扰,它在这里表明,低频听觉皮层活动是可靠的同步到缓慢的时间调制的语音,即使当噪声是两倍的语音强度。这种可靠的神经表示是通过强度对比度增益控制,并通过在不同的时间尺度,对应于神经δ和θ带的时间调制的自适应处理。重要的是,这种神经同步的精确性预测了听者在噪声中识别语音的能力,这表明听觉皮层表征的精确性限制了噪声中语音识别的性能。两者合计,这些结果表明,在一个复杂的听力环境中,听觉皮层可以选择性地编码语音流在一个背景不敏感的方式,这种稳定的语音神经表示提供了一个合理的基础背景不变的语音识别。
Speech recognition is remarkably robust to the listening background, even when the energy of background sounds strongly overlaps with that of speech. How the brain transforms the corrupted acoustic signal into a reliable neural representation suitable for speech recognition, however, remains elusive. Here, we hypothesize that this transformation is performed at the level of auditory cortex through adaptive neural encoding, and we test the hypothesis by recording, using magnetoencephalography (MEG), the neural responses of human subjects listening to a narrated story. Spectrally matched stationary noise, which has maximal acoustic overlap with the speech, is mixed in at various intensity levels. Despite the severe acoustic interference caused by this noise, it is here demonstrated that low-frequency auditory cortical activity is reliably synchronized to the slow temporal modulations of speech, even when the noise is twice as strong as the speech. Such a reliable neural representation is maintained by intensity contrast gain control, and by adaptive processing of temporal modulations at different time scales, corresponding to the neural delta and theta bands. Critically, the precision of this neural synchronization predicts how well a listener can recognize speech in noise, indicating that the precision of the auditory cortical representation limits the performance of speech recognition in noise. Taken together, these results suggest that, in a complex listening environment, auditory cortex can selectively encode a speech stream in a background insensitive manner, and this stable neural representation of speech provides a plausible basis for background-invariant recognition of speech.