Temporal coherence and the streaming of complex sounds.

Temporal coherence and the streaming of complex sounds.
复制标题

时间连贯性和复杂声音的流动。

DOI:
10.1007/978-1-4614-1590-9_59
复制
发表时间:
2013
影响因子:
--
通讯作者:
Xu,Yanbo
Xu,Yanbo
中科院分区:
医学4区
文献类型:
--
作者:
Shamma,Shihab;Elhilali,Mounya;Ma,Ling;Micheyl,Christophe;Oxenham,AndrewJ;Pressnitzer,Daniel;Yin,Pingbo;Xu,Yanbo

文献摘要

被引文献

相似文献

人类和其他动物可以注意到多种声音中的一种,并随着时间的推移有选择地跟随它。这种感知能力的神经基础仍然是个谜。一些研究已经得出结论,当声音激活分离良好的中央听觉神经元时,它们被作为单独的流听到,并且这个过程在很大程度上是预先注意的。在这里,我们建议,而不是流的形成主要取决于时间的响应,编码的声源的各种功能之间的相干性。此外,我们假设,只有当注意力指向一个特定的特征(例如,间距或位置)对该源的所有其它时间相干特征(例如,音色和位置)结合在一起成为一个流,与其他来源的不相干特征分离。支持这一假设的实验神经生理学证据。然而,重点将放在这一想法的计算实现上,并讨论从模拟中学到的见解,以解开复杂的声源,如语音和音乐。该模型由早期和皮层听觉处理的代表性阶段组成,该阶段创建了对各种声音属性(如音高、位置和频谱分辨率)的多维描述。接下来的阶段计算一个相干矩阵,该矩阵总结了构成皮层表示的所有通道之间的成对相关性。最后,通过将相干矩阵分解成其不相关的分量来提取感知到的分离流。该模型提出的问题进行了讨论,特别是在流和搜索进一步的神经相关的流感知的注意力的作用。
Humans and other animals can attend to one of multiple sounds, and follow it selectively over time. The neural underpinnings of this perceptual feat remain mysterious. Some studies have concluded that sounds are heard as separate streams when they activate well-separated populations of central auditory neurons, and that this process is largely pre-attentive. Here, we propose instead that stream formation depends primarily on temporal coherence between responses that encode various features of a sound source. Furthermore, we postulate that only when attention is directed toward a particular feature (e.g., pitch or location) do all other temporally coherent features of that source (e.g., timbre and location) become bound together as a stream that is segregated from the incoherent features of other sources. Experimental neurophysiological evidence in support of this hypothesis will be presented. The focus, however, will be on a computational realization of this idea and a discussion of the insights learned from simulations to disentangle complex sound sources such as speech and music. The model consists of a representational stage of early and cortical auditory processing that creates a multidimensional depiction of various sound attributes such as pitch, location, and spectral resolution. The following stage computes a coherence matrix that summarizes the pair-wise correlations between all channels making up the cortical representation. Finally, the perceived segregated streams are extracted by decomposing the coherence matrix into its uncorrelated components. Questions raised by the model are discussed, especially on the role of attention in streaming and the search for further neural correlates of streaming percepts.