TextonBoost for Image Understanding: Multi-Class Object Recognition and Segmentation by Jointly Modeling Texture, Layout, and Context

TextonBoost for Image Understanding: Multi-Class Object Recognition and Segmentation by Jointly Modeling Texture, Layout, and Context
复制标题

DOI:
10.1007/s11263-007-0109-1
复制
发表时间:
2009-01-01
影响因子:
19.5
通讯作者:
Criminisi, Antonio
Criminisi, Antonio
中科院分区:
计算机科学2区
文献类型:
--
作者:
Shotton, Jamie;Winn, John;Criminisi, Antonio

文献摘要

被引文献

相似文献

本文详细介绍了一种新的方法来学习对象类的判别模型,有效地结合纹理,布局和上下文信息。学习模型用于自动视觉理解和语义分割的照片。我们的判别模型利用纹理布局过滤器,新的功能的基础上textons,共同模式的纹理和空间布局。一元分类和特征选择是通过共享提升来实现的,以给出一个可以应用于大量类的有效分类器。通过将一元分类器结合到条件随机场中来实现精确的图像分割,该条件随机场(i)捕获相邻像素的类标签之间的空间相互作用,以及(ii)改善特定对象实例的分割。通过利用随机特征选择和分段训练方法,在大型数据集上实现了模型的有效训练。在四个不同的数据库上证明了高分类和分割精度:(i)MSRC 21级数据库,包含在一般照明条件、姿态和视点下观察的真实的物体的照片,(ii)7类Corel子集和(iii)He等人(Proceeding of IEEE Conference on Computer Vision and Pattern Recognition,vol.2,pp. 695-702,2004年6月),和(iv)一组电视节目的视频序列。所提出的算法为高度纹理化的对象(草、树等)提供了有竞争力的和视觉上令人愉悦的结果,高度结构化(汽车、人脸、自行车、飞机等),甚至是铰接的(身体、奶牛等)。
This paper details a new approach for learning a discriminative model of object classes, incorporating texture, layout, and context information efficiently. The learned model is used for automatic visual understanding and semantic segmentation of photographs. Our discriminative model exploits texture-layout filters, novel features based on textons, which jointly model patterns of texture and their spatial layout. Unary classification and feature selection is achieved using shared boosting to give an efficient classifier which can be applied to a large number of classes. Accurate image segmentation is achieved by incorporating the unary classifier in a conditional random field, which (i) captures the spatial interactions between class labels of neighboring pixels, and (ii) improves the segmentation of specific object instances. Efficient training of the model on large datasets is achieved by exploiting both random feature selection and piecewise training methods.High classification and segmentation accuracy is demonstrated on four varied databases: (i) the MSRC 21-class database containing photographs of real objects viewed under general lighting conditions, poses and viewpoints, (ii) the 7-class Corel subset and (iii) the 7-class Sowerby database used in He et al. (Proceeding of IEEE Conference on Computer Vision and Pattern Recognition, vol. 2, pp. 695-702, June 2004), and (iv) a set of video sequences of television shows. The proposed algorithm gives competitive and visually pleasing results for objects that are highly textured (grass, trees, etc.), highly structured (cars, faces, bicycles, airplanes, etc.), and even articulated (body, cow, etc.).