Learning a dictionary of shape-components in visual cortex: comparison with neurons, humans and machines

Learning a dictionary of shape-components in visual cortex: comparison with neurons, humans and machines
复制标题

DOI:
--
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
T. Poggio;T. Serre
T. Poggio;T. Serre
中科院分区:
其他
文献类型:
--
作者:
T. Poggio;T. Serre

文献摘要

被引文献

相似文献

在本文中,我描述了一个定量模型,该模型解释了视觉皮层腹侧流的前馈路径的电路和计算。该模型与视觉处理的一般理论一致,该理论将 [Hubel 和 Wiesel,1959] 的层次模型从初级视觉区域扩展到了纹状体视觉区域。它试图解释最初的几百毫秒的视觉处理和“立即识别”。该方法的关键要素之一是学习从 V2 到 IT 的形状组件通用字典,它为高级大脑区域中的任务特定分类电路提供了不变的表示。这种形状调整单元的词汇是以无监督的方式从自然图像中学习的,并且构成了具有不同复杂性和不变性的大量冗余图像特征集。该理论显着扩展了 [Riesenhuber 和 Poggio, 1999a] 的早期方法,并建立在几个现有的神经生物学模型和概念建议的基础上。首先,我提供证据表明该模型可以复制不同大脑区域(例如 V1、V4 和 IT)神经元的调节特性。特别是,该模型与 V4 中关于神经元对简单两杆刺激组合的反应的数据一致 [Reynolds et al., 1999](在 S2 单元的感受野内),并且模型中的一些 C2 单元显示出边界构象的调整,这与 V4 的记录一致 [Pasupathy 和 Connor, 2001]。其次,我证明该模型不仅可以在人工刺激下复制不同大脑区域神经元的调节特性,而且还可以处理现实世界中物体的识别,达到与最好的计算机视觉系统竞争的程度。第三,我描述了模型的性能与人类观察者在快速动物与非动物识别任务中的性能之间的比较,其中识别速度很快,而皮质反投影可能不活跃。结果表明,当刺激和面罩之间的延迟约为 50 毫秒时,该模型可以非常好地预测人类表现。这表明当时间间隔在此范围内时,皮质反投影可能不会发挥重要作用,因此该模型可以提供前馈路径的令人满意的描述。总而言之,证据表明我们可能拥有成功的视觉皮层理论的骨架。此外,这可能是第一次,忠实于视觉皮层生理学和解剖学的神经生物学模型不仅可以与一些最好的计算机视觉系统竞争,从而为工程人工视觉系统提供现实的替代方案,而且在涉及复杂自然图像的分类任务中实现接近人类的性能。 (副本可从麻省理工学院图书馆独家获取,地址:Rm. 14-0551, Cambridge, MA 02139-4307。电话:617-253-5668;传真:617-253-1690。)
In this thesis, I describe a quantitative model that accounts for the circuits and computations of the feedforward path of the ventral stream of visual cortex. This model is consistent with a general theory of visual processing that extends the hierarchical model of [Hubel and Wiesel, 1959] from primary to extrastriate visual areas. It attempts to explain the first few hundred milliseconds of visual processing and "immediate recognition". One of the key elements in the approach is the learning of a generic dictionary of shape-components from V2 to IT, which provides an invariant representation to task-specific categorization circuits in higher brain areas. This vocabulary of shape-tuned units is learned in an unsupervised manner from natural images, and constitutes a large and redundant set of image features with different complexities and invariances. This theory significantly extends an earlier approach by [Riesenhuber and Poggio, 1999a] and builds upon several existing neurobiological models and conceptual proposals. First, I present evidence to show that the model can duplicate the tuning properties of neurons in various brain areas (e.g., V1, V4 and IT). In particular, the model agrees with data from V4 about the response of neurons to combinations of simple two-bar stimuli [Reynolds et al., 1999] (within the receptive field of the S2 units) and some of the C2 units in the model show a tuning for boundary conformations which is consistent with recordings from V4 [Pasupathy and Connor, 2001]. Second, I show that not only can the model duplicate the tuning properties of neurons in various brain areas when probed with artificial stimuli, but it can also handle the recognition of objects in the real-world, to the extent of competing with the best computer vision systems. Third, I describe a comparison between the performance of the model and the performance of human observers in a rapid animal vs. non-animal recognition task for which recognition is fast and cortical back-projections are likely to be inactive. Results indicate that the model predicts human performance extremely well when the delay between the stimulus and the mask is about 50 ms. This suggests that cortical back-projections may not play a significant role when the time interval is in this range, and the model may therefore provide a satisfactory description of the feedforward path. Taken together, the evidences suggest that we may have the skeleton of a successful theory of visual cortex. In addition, this may be the first time that a neurobiological model, faithful to the physiology and the anatomy of visual cortex, not only competes with some of the best computer vision systems thus providing a realistic alternative to engineered artificial vision systems, but also achieves performance close to that of humans in a categorization task involving complex natural images. (Copies available exclusively from MIT Libraries, Rm. 14-0551, Cambridge, MA 02139-4307. Ph. 617-253-5668; Fax 617-253-1690.)