Support vector description of clusters for content-based image annotation

Support vector description of clusters for content-based image annotation
复制标题

DOI:
10.1016/j.patcog.2013.10.015
复制
发表时间:
2014-03
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Liang Sun;H. Ge;Shinichi Yoshida;Yanchun Liang;Guozhen Tan
Liang Sun;H. Ge;Shinichi Yoshida;Yanchun Liang;Guozhen Tan
中科院分区:
其他
文献类型:
--
作者:
Liang Sun;H. Ge;Shinichi Yoshida;Yanchun Liang;Guozhen Tan

文献摘要

被引文献

相似文献

计算机视觉和机器学习领域的不断进步为开发用于标记图像的自动工具提供了机会;这有助于搜索和检索。然而,由于现实世界的图像系统的复杂性,有效和高效的图像标注仍然是一个具有挑战性的问题。在本文中,我们提出了一个注释技术的基础上使用的图像内容和文字的相关性。具有手动标记的单词的图像集群被用作训练实例。使用核方法对每个聚类内的图像进行建模,其中图像向量被映射到更高维的空间,并且被识别为支持向量的向量被用于描述聚类。为了测量图像和由支持向量描述的模型之间的关联程度,计算从图像到模型的距离。距离越近,关联性越强。此外,词与词的相关性也被考虑在注释框架。为了标记图像,系统通过使用从图像到模型的距离和词与词的相关性在统一的概率框架中预测注释词。在三个基准图像数据集上进行了模拟实验。结果表明,所提出的技术的性能,并将其与其他最近报道的技术的性能进行比较。
Continual progress in the fields of computer vision and machine learning has provided opportunities to develop automatic tools for tagging images; this facilitates searching and retrieving. However, due to the complexity of real-world image systems, effective and efficient image annotation is still a challenging problem. In this paper, we present an annotation technique based on the use of image content and word correlations. Clusters of images with manually tagged words are used as training instances. Images within each cluster are modeled using a kernel method, in which the image vectors are mapped to a higher-dimensional space and the vectors identified as support vectors are used to describe the cluster. To measure the extent of the association between an image and a model described by support vectors, the distance from the image to the model is computed. A closer distance indicates a stronger association. Moreover, word-to-word correlations are also considered in the annotation framework. To tag an image, the system predicts the annotation words by using the distances from the image to the models and the word-to-word correlations in a unified probabilistic framework. Simulated experiments were conducted on three benchmark image data sets. The results demonstrate the performance of the proposed technique, and compare it to the performance of other recently reported techniques.