A Hierarchical and Contextual Model for Aerial Image Parsing

A Hierarchical and Contextual Model for Aerial Image Parsing
复制标题

DOI:
10.1007/s11263-009-0306-1
复制
发表时间:
2010-06
影响因子:
19.5
通讯作者:
Jake Porway;Qiongchen Wang;Song-Chun Zhu
Jake Porway;Qiongchen Wang;Song-Chun Zhu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jake Porway;Qiongchen Wang;Song-Chun Zhu

文献摘要

被引文献

相似文献

在本文中,我们提出了一个层次和上下文模型的航空图像理解。我们的模型将空中场景中的对象(汽车、屋顶、道路、树木、停车场)组织成分层组,其外观和配置由统计约束(例如,相对位置、相对尺度等)确定。我们的层次结构是一个非递归的语法,在航空图像中的对象组成的层的节点,每个可以分解成一些不同的配置。这使我们能够用相对较少的规则生成和识别大量的场景。我们提出了一个极小极大熵框架学习对象之间的统计约束,并表明,这种学习的背景下,使我们能够排除不太可能的场景配置和幻觉未检测到的对象在推理。类似的算法被提出用于纹理合成(Zhu et al.在Int. J. Comput中。维斯。2:107-126,1998),但是没有并入分层信息。我们使用一系列不同的自下而上检测器(AdaBoost,TextonBoost,Compositional Boosting(Freund和Schapire in J. Comput.系统科学55,1997;肖顿等人在Proceedings of the European Conference on Computer Vision,第1-15页,2006年; Wu et al. IEEE计算机视觉与模式识别会议论文集,pp. 1-8,2007))来提出对象在新的航空图像中的位置,并采用聚类采样算法(C4(Porway和Zhu,2009))来选择根据我们学习的先验模型最好地解释图像的检测子集。C4算法可以快速有效地在备选的竞争子解之间切换,例如,图像块是否由具有汽车的停车场或具有通风口的建筑物更好地解释。我们还表明,我们的模型可以预测我们的探测器错过的对象的位置。最后,我们提出解析的航拍图像和实验结果表明,我们的聚类采样和自上而下的预测算法使用从我们的模型中学习的上下文线索,以提高检测结果比传统的自下而上的检测器单独。
In this paper we present a hierarchical and contextual model for aerial image understanding. Our model organizes objects (cars, roofs, roads, trees, parking lots) in aerial scenes into hierarchical groups whose appearances and configurations are determined by statistical constraints (e.g. relative position, relative scale, etc.). Our hierarchy is a non-recursive grammar for objects in aerial images comprised of layers of nodes that can each decompose into a number of different configurations. This allows us to generate and recognize a vast number of scenes with relatively few rules. We present a minimax entropy framework for learning the statistical constraints between objects and show that this learned context allows us to rule out unlikely scene configurations and hallucinate undetected objects during inference. A similar algorithm was proposed for texture synthesis (Zhu et al. in Int. J. Comput. Vis. 2:107–126, 1998) but didn’t incorporate hierarchical information. We use a range of different bottom-up detectors (AdaBoost, TextonBoost, Compositional Boosting (Freund and Schapire in J. Comput. Syst. Sci. 55, 1997; Shotton et al. in Proceedings of the European Conference on Computer Vision, pp. 1–15, 2006; Wu et al. in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1–8, 2007)) to propose locations of objects in new aerial images and employ a cluster sampling algorithm (C4 (Porway and Zhu, 2009)) to choose the subset of detections that best explains the image according to our learned prior model. The C4 algorithm can quickly and efficiently switch between alternate competing sub-solutions, for example whether an image patch is better explained by a parking lot with cars or by a building with vents. We also show that our model can predict the locations of objects our detectors missed. We conclude by presenting parsed aerial images and experimental results showing that our cluster sampling and top-down prediction algorithms use the learned contextual cues from our model to improve detection results over traditional bottom-up detectors alone.