METRIC INVARIANCE IN OBJECT RECOGNITION - A REVIEW AND FURTHER EVIDENCE

METRIC INVARIANCE IN OBJECT RECOGNITION - A REVIEW AND FURTHER EVIDENCE
复制标题

DOI:
10.1037/h0084317
复制
发表时间:
1992-06-01
期刊:
CANADIAN JOURNAL OF PSYCHOLOGY-REVUE CANADIENNE DE PSYCHOLOGIE
影响因子:
--
通讯作者:
HUMMEL, JE
HUMMEL, JE
中科院分区:
其他
文献类型:
--
作者:
COOPER, EE;BIEDERMAN, I;HUMMEL, JE

文献摘要

被引文献

相似文献

从现象学的角度来看,人类的形状识别似乎是不变的,随着深度方向的变化(直到部分遮挡),在视野中的位置和大小的变化。最近版本的模板理论(例如,Ullman, 1989; Lowe, 1987)假设这些不变性是通过应用图像的旋转、平移和缩放等变换来实现的,这样它就可以与存储的模板进行度量匹配。大概,这样的转换需要时间来执行。我们描述了最近的启动实验,其中对图像的先前简短呈现对其后续识别的影响进行了评估。这些实验的结果表明,不变性是完全的:视觉启动的大小(不同于名称或基本级别的概念启动)不受位置、大小、深度方向或图像中存在的特定直线和顶点的变化的影响,只要相同组件的表示可以被激活。描述了一个实现的七层神经网络模型(Hummel & Biederman, 1992),该模型捕获了人类物体识别的这些基本属性。给定对象的线条图,模型激活对象的视点不变结构描述,指定其部分及其相互关系。视觉启动被解释为激活的连接权重的变化:a)被称为geon feature assemblies (GFAS)的单元,它将代表单个geon及其关系(如其类型,长宽比,与其他geon的关系)的不变,独立属性的单元输出连接在一起,或者b)连接权重的变化,通过几个gfa激活代表对象的cell。
Phenomenologically, human shape recognition appears to be invariant with changes of orientation in depth (up to parts occlusion), position in the visual field, and size. Recent versions of template theories (e.g., Ullman, I989; Lowe, I987) assume that these invariances are achieved through the application of transformations such as rotation, translation, and scaling of the image so that it can be matched metrically to a stored template. Presumably, such transformations would require time for their execution. We describe recent priming experiments in which the effects of a prior brief presentation of an image on its subsequent recognition are assessed. The results of these experiments indicate that the invariance is complete: The magnitude of visual priming (as distinct from name or basic level concept priming) is not affected by a change in position, size, orientation in depth, or the particular lines and vertices present in the image, as long as representations of the same components can be activated. An implemented seven layer neural network model (Hummel & Biederman, I992) that captures these fundamental properties of human object recognition is described. Given a line drawing of an object, the model activates a viewpoint-invariant structural description of the object, specifying its parts and their interrelations. Visual priming is interpreted as a change in the connection weights for the activation of: a) cells, termed geon feature assemblies (GFAS), that conjoin the output of units that represent invariant, independent properties of a single geon and its relations (such as its type, aspect ratio, relations to other geons), or b) a change in the connection weights by which several GFAs activate a cell representing an object.