Noise-robust cortical tracking of attended speech in real-world acoustic scenes

Noise-robust cortical tracking of attended speech in real-world acoustic scenes
复制标题

DOI:
10.1016/j.neuroimage.2017.04.026
复制
发表时间:
2017-08-01
期刊:
影响因子:
5.7
通讯作者:
Hjortkjaer, Jens
Hjortkjaer, Jens
中科院分区:
医学1区
文献类型:
--
作者:
Fuglsang, Soren Asp;Dau, Torsten;Hjortkjaer, Jens

文献摘要

被引文献

相似文献

在多说话者场景中选择性地关注一个说话者被认为是将低频皮层活动与所关注的语音信号同步。在最近的研究中,从单次试验脑电图(EEG)数据的语音重建已被用于解码在两个说话者的情况下,听者正在关注哪个说话者。目前尚不清楚这如何推广到更复杂的声音环境。从行为上讲,语音感知对听众在日常生活中通常遇到的声学失真是鲁棒的,但不知道这是否反映了对语音的噪声鲁棒神经跟踪。在这里,我们使用先进的声学模拟在实验室中重现真实世界的声学场景。在具有不同混响量和干扰说话者数量的虚拟声学现实中,听众选择性地关注特定说话者的语音流。在不同的听力环境中,我们发现,参加谈话者可以准确地从单次试验EEG数据解码,而不管在声学输入的不同失真。对于高度混响的环境中,语音包络重建神经反应失真的刺激更像原来的干净的信号比失真的输入。与混响语音,我们观察到一个迟到的皮质响应出席的语音流编码的时间调制的语音信号没有其混响失真。基于来自64个头皮电极的40-50 s长数据块的单次试验注意力解码准确率在所有考虑的听力环境中都同样高(80-90%正确),并且在使用10个头皮电极和短(
Selectively attending to one speaker in a multi-speaker scenario is thought to synchronize low-frequency cortical activity to the attended speech signal. In recent studies, reconstruction of speech from single-trial electroencephalogram (EEG) data has been used to decode which talker a listener is attending to in a two-talker situation. It is currently unclear how this generalizes to more complex sound environments. Behaviorally, speech perception is robust to the acoustic distortions that listeners typically encounter in everyday life, but it is unknown whether this is mirrored by a noise-robust neural tracking of attended speech. Here we used advanced acoustic simulations to recreate real-world acoustic scenes in the laboratory. In virtual acoustic realities with varying amounts of reverberation and number of interfering talkers, listeners selectively attended to the speech stream of a particular talker. Across the different listening environments, we found that the attended talker could be accurately decoded from single-trial EEG data irrespective of the different distortions in the acoustic input. For highly reverberant environments, speech envelopes reconstructed from neural responses to the distorted stimuli resembled the original clean signal more than the distorted input. With reverberant speech, we observed a late cortical response to the attended speech stream that encoded temporal modulations in the speech signal without its reverberant distortion. Single-trial attention decoding accuracies based on 40-50 s long blocks of data from 64 scalp electrodes were equally high (80-90% correct) in all considered listening environments and remained statistically significant using down to 10 scalp electrodes and short (