Disentangling the Independent Contributions of Visual and Conceptual Features to the Spatiotemporal Dynamics of Scene Categorization

Disentangling the Independent Contributions of Visual and Conceptual Features to the Spatiotemporal Dynamics of Scene Categorization
复制标题

DOI:
10.1523/jneurosci.2088-19.2020
复制
发表时间:
2020-07-01
影响因子:
5.3
通讯作者:
Hansen, Bruce C.
Hansen, Bruce C.
中科院分区:
医学1区
文献类型:
--
作者:
Greene, Michelle R.;Hansen, Bruce C.

文献摘要

被引文献

相似文献

人类场景分类具有速度快的特点。虽然许多视觉和概念特征与这种能力有关,但特征空间之间存在显着的相关性,阻碍了我们确定其对场景分类的相对贡献的能力。在这里,我们使用了白化变换去相关的各种视觉和概念的功能,并评估其独特的贡献场景分类的时间过程。参与者(男性和女性)观看了2250张从30个不同场景类别中绘制的全彩色场景图像,同时通过256通道EEG测量了他们的大脑活动。我们研究了解释在每个电极和时间点的视觉事件相关电位(vERP)数据从9个不同的白化编码模型的方差。这些范围从低层次的功能,从过滤器输出到高层次的概念特征,需要人类注释。通过多变量解码方法评估vERPs中的类别信息量。在单独的众包实验中获得行为相似性度量。我们发现,所有9个模型共同贡献了78%的人类场景相似性评估的方差,并且在vERP数据的噪声上限内。低水平模型解释了早期vERP的变异性(图像开始后88 ms),而高水平模型解释了后来的方差(169 ms)。重要的是,只有高级模型与行为共享vERP可变性。总之,这些结果表明,场景分类主要是一个高层次的过程,但依赖于以前提取的低层次的功能。
Human scene categorization is characterized by its remarkable speed. While many visual and conceptual features have been linked to this ability, significant correlations exist between feature spaces, impeding our ability to determine their relative contributions to scene categorization. Here, we used a whitening transformation to decorrelate a variety of visual and conceptual features and assess the time course of their unique contributions to scene categorization. Participants (both sexes) viewed 2250 full colorscene images drawn from 30 different scene categories while having their brain activity measured through 256-channel EEG. We examined the variance explained at each electrode and time point of visual event related potential (vERP) data from nine different whitened encoding models. These ranged from low-level features obtained from filter outputs to high level conceptual features requiring human annotation. The amount of category information in the vERPs was assessed through multivariate decoding methods. Behavioral similarity measures were obtained in separate crowdsourced experiments. We found that all nine models together contributed 78% of the variance of human scene similarity assessments and were within the noise ceiling of the vERP data. Low-level models explained earlier vERP variability (88 ms after image onset), whereas high-level models explained later variance (169 ms). Critically, only high-level models shared vERP variability with behavior. Together, these results suggest that scene categorization is primarily a high-level process, but reliant on previously extracted low-level features.