Modeling the shape of the scene: A holistic representation of the spatial envelope

Modeling the shape of the scene: A holistic representation of the spatial envelope
复制标题

DOI:
10.1023/a:1011139631724
复制
发表时间:
2001-01-01
影响因子:
19.5
通讯作者:
Torralba, A
Torralba, A
中科院分区:
计算机科学2区
文献类型:
--
作者:
Oliva, A;Torralba, A

文献摘要

被引文献

相似文献

在本文中,我们提出了一种绕过单个对象或区域的分割和处理的真实场景识别的计算模型。该过程基于场景的非常低维的表示,我们称之为空间包络。我们提出了一组感知维度(自然度、开放性、粗糙度、扩张性、粗糙度)来表示场景的主导空间结构。然后,我们证明了使用谱和粗局域化信息可以可靠地估计这些维度。该模型生成多维空间,其中共享语义类别(例如,街道、高速公路、海岸)中的成员资格的场景被紧密地投影在一起。空间包络模型的性能表明,关于对象形状或身份的特定信息不是场景分类的要求,并且对场景的整体表示进行建模可以告知其可能的语义类别。
In this paper, we propose a computational model of the recognition of real world scenes that bypasses the segmentation and the processing of individual objects or regions. The procedure is based on a very low dimensional representation of the scene, that we term the Spatial Envelope. We propose a set of perceptual dimensions (naturalness, openness, roughness, expansion, ruggedness) that represent the dominant spatial structure of a scene. Then, we show that these dimensions may be reliably estimated using spectral and coarsely localized information. The model generates a multidimensional space in which scenes sharing membership in semantic categories (e.g., streets, highways, coasts) are projected closed together. The performance of the spatial envelope model shows that specific information about object shape or identity is not a requirement for scene categorization and that modeling a holistic representation of the scene informs about its probable semantic category.