Beyond the Euclidean distance: Creating effective visual codebooks using the Histogram Intersection Kernel

Beyond the Euclidean distance: Creating effective visual codebooks using the Histogram Intersection Kernel
复制标题

DOI:
10.1109/iccv.2009.5459178
复制
发表时间:
2009-09
期刊:
2009 IEEE 12th International Conference on Computer Vision
影响因子:
--
通讯作者:
Jianxin Wu;James M. Rehg
Jianxin Wu;James M. Rehg
中科院分区:
其他
文献类型:
--
作者:
Jianxin Wu;James M. Rehg

文献摘要

被引文献

相似文献

Bag of Visual Words 模型中使用的常见视觉码本生成方法,例如k-means 或高斯混合模型,使用欧几里得距离将特征聚类为视觉码字。然而,最流行的视觉描述符是图像测量的直方图。事实证明,在具有直方图特征的监督学习任务中,直方图交集核(HIK)比欧几里得距离更有效。在本文中,我们证明 HIK 也可以以无监督的方式使用,以显着改进视觉码本的生成。我们提出了一种直方图内核 k-means 算法,该算法易于实现并且运行速度几乎与 k-means 一样快。 HIK 码本的识别准确率始终比 k-means 码本高 2-4%。此外,我们提出了一种一类 SVM 公式来创建更有效的视觉码字,从而实现更高的准确性。所提出的方法为 3 个流行的对象和场景识别基准数据集建立了新的最先进的性能数据。此外,我们还表明标准 k 中值聚类方法可用于视觉码本生成,并且可以作为 HIK 和 k 均值方法之间的折衷方案。
Common visual codebook generation methods used in a Bag of Visual words model, e.g. k-means or Gaussian Mixture Model, use the Euclidean distance to cluster features into visual code words. However, most popular visual descriptors are histograms of image measurements. It has been shown that the Histogram Intersection Kernel (HIK) is more effective than the Euclidean distance in supervised learning tasks with histogram features. In this paper, we demonstrate that HIK can also be used in an unsupervised manner to significantly improve the generation of visual codebooks. We propose a histogram kernel k-means algorithm which is easy to implement and runs almost as fast as k-means. The HIK codebook has consistently higher recognition accuracy over k-means codebooks by 2–4%. In addition, we propose a one-class SVM formulation to create more effective visual code words which can achieve even higher accuracy. The proposed method has established new state-of-the-art performance numbers for 3 popular benchmark datasets on object and scene recognition. In addition, we show that the standard k-median clustering method can be used for visual codebook generation and can act as a compromise between HIK and k-means approaches.