Linear spatial pyramid matching using sparse coding for image classification

Linear spatial pyramid matching using sparse coding for image classification
复制标题

DOI:
10.1109/cvpr.2009.5206757
复制
发表时间:
2009-06
期刊:
2009 IEEE Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Jianchao Yang;Kai Yu;Yihong Gong;Thomas Huang
Jianchao Yang;Kai Yu;Yihong Gong;Thomas Huang
中科院分区:
其他
文献类型:
--
作者:
Jianchao Yang;Kai Yu;Yihong Gong;Thomas Huang

文献摘要

被引文献

相似文献

近年来,基于空间金字塔匹配(SPM)核的支持向量机在图像分类中取得了很大的成功。尽管它很受欢迎,但这些非线性支持向量机的训练复杂度为O(n2 ~ n3),测试复杂度为O(n),其中n是训练大小,这意味着将算法扩展到处理数千个训练图像是不平凡的。在本文中,我们开发了一个扩展的SPM方法,通过推广矢量量化的稀疏编码,其次是多尺度空间最大池,并提出了一个线性SPM内核的基础上SIFT稀疏码。这种新的方法显着降低了支持向量机的复杂度为O(n)的训练和一个常数的测试。在一些图像分类实验中,我们发现,在分类精度方面,建议的线性SPM的基础上稀疏编码的SIFT描述符总是显着优于线性SPM内核的直方图,甚至优于非线性SPM内核,导致国家的最先进的性能在几个基准测试通过使用单一类型的描述符。
Recently SVMs using spatial pyramid matching (SPM) kernel have been highly successful in image classification. Despite its popularity, these nonlinear SVMs have a complexity O(n2 ~ n3) in training and O(n) in testing, where n is the training size, implying that it is nontrivial to scaleup the algorithms to handle more than thousands of training images. In this paper we develop an extension of the SPM method, by generalizing vector quantization to sparse coding followed by multi-scale spatial max pooling, and propose a linear SPM kernel based on SIFT sparse codes. This new approach remarkably reduces the complexity of SVMs to O(n) in training and a constant in testing. In a number of image categorization experiments, we find that, in terms of classification accuracy, the suggested linear SPM based on sparse coding of SIFT descriptors always significantly outperforms the linear SPM kernel on histograms, and is even better than the nonlinear SPM kernels, leading to state-of-the-art performance on several benchmarks by using a single type of descriptors.