Ultra-rapid object categorization in real-world scenes with top-down manipulations

Ultra-rapid object categorization in real-world scenes with top-down manipulations
复制标题

DOI:
10.1371/journal.pone.0214444
复制
发表时间:
2019-04-10
期刊:
影响因子:
3.7
通讯作者:
Zhao, Qi
Zhao, Qi
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Xu, Bingjie;Kankanhalli, Mohan S.;Zhao, Qi

文献摘要

被引文献

相似文献

人类能够快速而毫不费力地实现视觉对象识别。人们普遍认为,对象分类是通过自下而上和自上而下的认知处理之间的相互作用来实现的。在刺激短暂出现并且响应时间有限的超快速分类场景中,假设前馈信息的第一扫描足以区分对象是否存在于场景中。然而,反馈/自上而下的处理是否以及如何在如此短暂的时间内参与仍然是一个悬而未决的问题。为此,在这里,我们想研究不同的自上而下的操作,如类别级别,类别类型和现实世界的大小,在超快速分类中如何相互作用。我们已经构建了一个数据集,包括现实世界的场景图像与目标对象显示大小的内置测量。基于这组图像,我们测量了人类受试者的超快速对象分类性能。标准的前馈计算模型表示场景功能和一个国家的最先进的目标检测模型进行辅助调查。结果表明,1)动物性(动物,车辆,食物),2)抽象水平(人,运动),3)现实世界的大小(四个目标大小水平)对超快速分类过程的影响。这对支持自上而下处理的参与产生了影响,当快速分类某些对象时,例如在细粒度级别上的体育。我们在人类与模型比较方面的工作也揭示了两者可能的合作和整合,这可能对实验和计算视觉研究都有意义。所有收集的图像和行为数据以及代码和模型都可以在https://osf.io/mqwjz/上公开获得。
Humans are able to achieve visual object recognition rapidly and effortlessly. Object categorization is commonly believed to be achieved by interaction between bottom-up and top-down cognitive processing. In the ultra-rapid categorization scenario where the stimuli appear briefly and response time is limited, it is assumed that a first sweep of feedforward information is sufficient to discriminate whether or not an object is present in a scene. However, whether and how feedback/top-down processing is involved in such a brief duration remains an open question. To this end, here, we would like to examine how different top-down manipulations, such as category level, category type and real-world size, interact in ultra-rapid categorization. We have constructed a dataset comprising real-world scene images with a built-in measurement of target object display size. Based on this set of images, we have measured ultra-rapid object categorization performance by human subjects. Standard feedforward computational models representing scene features and a state-of-the-art object detection model were employed for auxiliary investigation. The results showed the influences from 1) animacy (animal, vehicle, food), 2) level of abstraction (people, sport), and 3) real-world size (four target size levels) on ultra-rapid categorization processes. This had an impact to support the involvement of top-down processing when rapidly categorizing certain objects, such as sport at a fine grained level. Our work on human vs. model comparisons also shed light on possible collaboration and integration of the two that may be of interest to both experimental and computational vision researches. All the collected images and behavioral data as well as code and models are publicly available at https://osf.io/mqwjz/.