Surprising twist on auditory representation. Focus on: "what's that sound? Auditory area CLM encodes stimulus surprise, not intensity or intensity changes".

Surprising twist on auditory representation. Focus on: "what's that sound? Auditory area CLM encodes stimulus surprise, not intensity or intensity changes".
复制标题

听觉表征的惊人转折。

DOI:
10.1152/jn.90270.2008
复制
发表时间:
2008
影响因子:
2.5
通讯作者:
Gentner,TimothyQ
Gentner,TimothyQ
中科院分区:
医学3区
文献类型:
--
作者:
Gentner,TimothyQ

文献摘要

相似文献

感觉表征的概念是感知和认知的心理学理论以及感觉编码的神经生物学模型的核心。然而,确切地理解神经反应“代表”的是一个很难完全回答的问题,特别是对于那些远离感觉传导点的神经元来说。Gill及其同事(2007)最近的一篇论文在理解听觉系统中的刺激表征方面迈出了一大步。为了理解本文的贡献,了解一些背景知识是有帮助的。第一种近似,感觉系统和它们所组成的感受野是按等级组织的。“感受野”的概念首先应用于视觉系统(Hartline 1938),以描述视神经纤维对照射在视网膜空间受限部分上的光的敏感性。虽然初级视觉皮层中的神经元继承了视网膜投射的空间感受野,但它们对定向的条形或边缘而不是光斑做出反应。在视觉通路的较高水平,细胞对复杂的视觉模式(如人脸或其他物体)反应最好(Suzuki et al. 2006; Tanaka 2003)。类似的表征层级存在于上升的听觉系统中,第八神经中初级纤维的频率调谐响应在更中央的水平上产生越来越复杂的“对象”(Gentner and Margoliash 2003)。尽管有现象学的描述,但初级听觉皮层以外的神经元的反应已经被证明非常难以建模。目前,最好的感受野模型通常占神经元对刺激的反应的25- 30%或更少(例如,Sen et al. 2001)。为了改进高级感觉神经元的感受野模型,人们可能会设计一个更好的函数来将刺激映射到尖峰串上-大多数模型依赖于线性回归(参见。Sharpee等人,2006年)。或者,正如Gill和他的同事所做的那样,人们可以设计一个更好的输入刺激表示(例如,Rust et al. 2006)。他们的研究结果表明,我们需要改变我们在听觉系统中思考时间的方式。当代高水平听觉神经反应的模型被称为“光谱-时间感受野”(或简称STRF),因为它们是根据每个尖峰之前的刺激的平均动态功率谱来定义的(Aertsen和Johannesma 1981)。最初的STRF概念和最近的变体(例如,Kowalski等人,1996年;泰尼森等人,2000年)对刺激的时间和光谱维度的处理方式类似。也就是说,刺激中跨时间播放的模式不会比跨频带播放的模式更大(或更小)。Gill等人为复杂的自然信号设计了一种新的表示方法,该方法考虑了刺激的展开时间结构。而不是考虑动态功率谱的大小(即,每个频率和每个时间点的功率)作为神经元的输入,他们将声谱图转换为一组大的条件概率,其中,在任何给定时间在每个频率处的值是在该频率处观察到的功率之前是在相邻频率处的特定功率模式的可能性的函数。想象一下,你坐在一个红绿灯前。红灯,然后突然变成绿色。如果以固定的时间间隔对交通信号灯的状态进行采样,则对于每对连续的时间间隔,存在三种可能的状态:在两个时间间隔中信号灯都是红色的,在两个时间间隔中信号灯都是绿色的,或者信号灯都是红色的。
The concept of sensory representation is central to psychological theories of perception and cognition and to neurobiological models of sensory coding. Understanding exactly what a neural response “represents,” however, turns out to be a difficult question to answer fully—particularly for neurons many synapses away from the point of sensory transduction. A recent paper by Gill and colleagues (2007) takes a large step forward in understanding stimulus representation in the auditory system. To appreciate the contribution this paper makes, it’s helpful to have a little background. To a first approximation, sensory systems and the receptive fields they comprise are organized hierarchically. The concept of a “receptive field” was first applied to the visual system (Hartline 1938) to describe the sensitivity of optic nerve fibers to light shined on spatially restricted portions of the retina. While neurons in primary visual cortex inherit the spatial receptive fields of the retinal projections, they respond to oriented bars or edges rather than spots of light. At higher levels in the visual pathway, cells respond best to complex visual patterns such as faces or other objects (Suzuki et al. 2006; Tanaka 2003). A similar representational hierarchy exists in the ascending auditory system, where the frequency tuned responses of primary fibers in the eighth nerve give rise to increasingly complex “objects” at more central levels (Gentner and Margoliash 2003). Phenomenological descriptions notwithstanding, the responses of neurons beyond primary auditory cortex have proven very difficult to model. Currently the best receptive field models typically account for 25–30%, or less, of a neuron’s response to a stimulus (eg, Sen et al. 2001). To improve the receptive field models for high-level sensory neurons, one might devise a better function to map the stimulus on to the spike train—most models rely on linear regression (cf. Sharpee et al. 2006). Alternatively, as Gill and colleagues did, one might devise a better representation of the input stimulus (eg, Rust et al. 2006). Their results suggest that we need to change the way we’ve been thinking about time in the auditory system. Contemporary models of high-level auditory neural responses are called “spectro-temporal receptive fields”(or STRFs for short) because they are defined in terms of the average dynamic power spectrum of the stimulus that precedes each spike (Aertsen and Johannesma 1981). The original STRF conception, and more recent variations (eg, Kowalski et al. 1996; Theunissen et al. 2000), treat the temporal and spectral dimensions of the stimulus similarly. That is, a pattern in the stimulus that plays out across time is given no greater (or lesser) weight than a pattern that plays out across frequency bands. Gill et al. devised a novel representation for complex natural signals that takes into account the unfolding temporal structure of the stimulus. Instead of considering the magnitude of the dynamic power spectrum (ie, the power at each frequency and each point in time) as the input to a neuron, they converted the sound spectrogram to a large set of conditional probabilities, where the value at each frequency at any given time is a function of the likelihood that the observed power at that frequency is preceded by a particular pattern of power at neighboring frequencies.To help think about this, imagine that you are sitting at a stop light. The light is red, the light is red, the light is red, and then suddenly it snaps to green. If one samples the state of the traffic light at regular time intervals, there are three possible states for each pair of sequential intervals: the light is red in both intervals, the light is green in both intervals, or the light is red …