View-based models of 3D object recognition: invariance to imaging transformations.

View-based models of 3D object recognition: invariance to imaging transformations.
复制标题

基于视图的 3D 对象识别模型:成像变换的不变性。

DOI:
10.1093/cercor/5.3.261
复制
发表时间:
1995
期刊:
Cerebral cortex (New York, N.Y. : 1991)
影响因子:
--
通讯作者:
Poggio,T
Poggio,T
中科院分区:
--
文献类型:
--
作者:
Vetter,T;Hurlbert,A;Poggio,T

文献摘要

被引文献

相似文献

本报告介绍了基于视图的对象识别模型的主要功能。该模型并不试图解释特定的皮层结构,它试图捕捉一般属性预期在生物结构的对象识别。基本模块是一个正则化网络(类似于RBF;参见Poggio和Girosi,1989; Poggio,1990),其中每个隐藏单元都广泛地调整到要识别的对象的特定视图。网络输出,这可能是在很大程度上独立于视图,首先描述了一些简单的模拟。然后讨论了基本模块的以下改进和细节:(1)某些单元可能只代表对象视图的组成部分--单元的最佳刺激(其“中心”)实际上是一个复杂的特征:(2)这些单元的性质与通常描述的皮层神经元调谐到多维最佳刺激一致,并且可以根据合理的生物物理机制来实现;(3)在学习识别新对象时,可以使用和修改预先存在的中心,但也可以递增地创建新中心,以便提供最大的视图不变性;(4)模块是分层结构的一部分-网络的输出可以用作另一个的输入之一,以这种方式合成越来越复杂的特征和模板;(5)在几个识别任务中,特别是在基本级别,使用视点不变特征的单个中心可能就足够了。这种类型的模块可以处理特定对象的识别,例如,在各种变换下的特定面部,诸如由于视点和照明引起的变换,只要特定对象的足够数量的示例视图可用。然而,3D对象识别的架构必须在某种程度上科普,即使只有一个单一的模型视图。这份报告的主要贡献是一个轮廓的识别架构,处理对象的一个很好的类经历了广泛的变换-由于照明,姿势,表情,等等-通过利用原型的例子。一个好的对象类是一组在特定变换(如视点变换)下具有足够相似的变换属性的对象。对于好的对象类,我们讨论了两种可能性:(1)类特定的转换将被应用于单个模型图像以生成附加的virtualexample视图,从而允许某种程度的泛化超出单个模型视图所能提供的范围;(2)从类的示例中学习类特定的、视图不变的特征,并将其与新的模型图像一起使用,而没有虚拟示例的显式生成。
This report describes the main features of a view-based model of object recognition. The model does not attempt to account for specific cortical structures; it tries to capture general properties to be expected in a biological architecture for object recognition. The basic module is a regularization network (RBF-like; see Poggio and Girosi, 1989; Poggio, 1990) in which each of the hidden units is broadly tuned to a specific view of the object to be recognized. The network output, which may be largely view independent, is first described in terms of some simple simulations. The following refinements and details of the basic module are then discussed: (1) some of the units may represent only components of views of the object—the optimal stimulus for the unit its “center,” is effectively a complex feature; (2) the units' properties are consistent with the usual description of cortical neurons as tuned to multidimensional optimal stimuli and may be realized in terms of plausible biophysical mechanisms; (3) in learning to recognize new objects, preexisting centers may be used and modified, but also new centers may be created incrementally so as to provide maximal view invariance; (4) modules are part of a hierarchical structure—the output of a network may be used as one of the inputs to another, in this way synthesizing increasingly complex features and templates; (5) in several recognition tasks, in particular at the basic level, a single center using view-invariant features may be sufficient.Modules of this type can deal with recognition of specific objects, for instance, a specific face under various transformations such as those due to viewpoint and illumination, provided that a sufficient number of example views of the specific object are available. An architecture for 3D object recognition, however, must cope- to some extent—even when only a single model view is given. The main contribution of this report is an outline of a recognition architecture that deals with objects of a nice class undergoing a broad spectrum of transformations—due to illumination, pose, expression, and so on- by exploiting prototypical examples. A nice class of objects is a set of objects with sufficiently similar transformation properties under specific transformations, such as viewpoint transformations. For nice object classes, we discuss two possibilities: (1) class-specific transformations are to be applied to a single model image to generate additionalvirtualexample views, thus allowing some degree of generalization beyond what a single model view could otherwise provide; (2) class-specific, view-invariant features are learned from examples of the class and used with the novel model image, without an explicit generation of virtual examples.