Fisher Kernels on Visual Vocabularies for Image Categorization

Fisher Kernels on Visual Vocabularies for Image Categorization
复制标题

DOI:
10.1109/cvpr.2007.383266
复制
发表时间:
2007-06
期刊:
2007 IEEE Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Florent Perronnin;C. Dance
Florent Perronnin;C. Dance
中科院分区:
其他
文献类型:
--
作者:
Florent Perronnin;C. Dance

文献摘要

被引文献

相似文献

在模式分类领域,Fisher内核是一个强大的框架,结合了生成和歧视方法的优势。这个想法是用源自生成概率模型得出的梯度向量的信号表征,然后将此表示形式馈送到区分分类器中。我们建议将此框架应用于图像分类,其中输入信号是图像,而基础生成模型是视觉词汇:一种高斯混合模型,该模型近似于图像中低级特征的分布。我们表明,Fisher内核实际上可以被理解为流行的维斯特袋的延伸。我们的方法在两个具有挑战性的数据库上表现出了出色的性能:一个内部数据库的19个对象/场景类别和最近发布的VOC 2006数据库。它也非常实用:在培训和测试时间和在一组类别中训练的词汇量都可以应用于另一组套件,而没有任何绩效损失。
Within the field of pattern classification, the Fisher kernel is a powerful framework which combines the strengths of generative and discriminative approaches. The idea is to characterize a signal with a gradient vector derived from a generative probability model and to subsequently feed this representation to a discriminative classifier. We propose to apply this framework to image categorization where the input signals are images and where the underlying generative model is a visual vocabulary: a Gaussian mixture model which approximates the distribution of low-level features in images. We show that Fisher kernels can actually be understood as an extension of the popular bag-of-visterms. Our approach demonstrates excellent performance on two challenging databases: an in-house database of 19 object/scene categories and the recently released VOC 2006 database. It is also very practical: it has low computational needs both at training and test time and vocabularies trained on one set of categories can be applied to another set without any significant loss in performance.