TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation

TagProp: Discriminative metric learning in nearest neighbor models for image auto-annotation
复制标题

DOI:
10.1109/iccv.2009.5459266
复制
发表时间:
2009-09
期刊:
2009 IEEE 12th International Conference on Computer Vision
影响因子:
--
通讯作者:
M. Guillaumin;Thomas Mensink;J. Verbeek;C. Schmid
M. Guillaumin;Thomas Mensink;J. Verbeek;C. Schmid
中科院分区:
其他
文献类型:
--
作者:
M. Guillaumin;Thomas Mensink;J. Verbeek;C. Schmid

文献摘要

被引文献

相似文献

图像自动标注是计算机视觉中一个重要的开放性问题。对于这项任务,我们提出了TagProp,一个有区别地训练的最近邻模型。使用加权最近邻模型预测测试图像的标签,以利用标记的训练图像。邻居权重基于邻居秩或距离。TagProp允许通过直接最大化训练集中标签预测的对数似然来集成度量学习。以这种方式,我们可以最佳地联合收割机的图像相似性度量的集合,覆盖图像内容的不同方面,如局部形状描述符,或全局颜色直方图。我们还引入了一个词特定的sigmoidal调制的加权邻居标签预测,以提高罕见的单词的召回。我们研究了模型不同变体的性能,并与现有工作进行比较。我们提出了三个具有挑战性的数据集的实验结果。在这三个方面,与当前最先进的技术相比,TagProp有了显著的改进。
Image auto-annotation is an important open problem in computer vision. For this task we propose TagProp, a discriminatively trained nearest neighbor model. Tags of test images are predicted using a weighted nearest-neighbor model to exploit labeled training images. Neighbor weights are based on neighbor rank or distance. TagProp allows the integration of metric learning by directly maximizing the log-likelihood of the tag predictions in the training set. In this manner, we can optimally combine a collection of image similarity metrics that cover different aspects of image content, such as local shape descriptors, or global color histograms. We also introduce a word specific sigmoidal modulation of the weighted neighbor tag predictions to boost the recall of rare words. We investigate the performance of different variants of our model and compare to existing work. We present experimental results for three challenging data sets. On all three, TagProp makes a marked improvement as compared to the current state-of-the-art.