Bridging the Semantic Gap Between Image Contents and Tags

Bridging the Semantic Gap Between Image Contents and Tags
复制标题

DOI:
10.1109/tmm.2010.2051360
复制
发表时间:
2010-08-01
影响因子:
7.3
通讯作者:
King, Irwin
King, Irwin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ma, Hao;Zhu, Jianke;King, Irwin

文献摘要

被引文献

相似文献

随着Web 2.0应用的飞速发展,标签被广泛用于描述Web上的图像内容。由于人类产生的标签具有噪声和稀疏性,如何理解和利用这些标签进行图像检索已成为一个新兴的研究方向。由于底层视觉特征能够提供丰富的信息,因此被用来提高图像检索的效果。然而,它是具有挑战性的图像内容和标签之间的语义鸿沟的桥梁。为了解决这个关键问题,本文提出了一个统一的框架,它源于图像内容和标签之间的两级数据融合:1)建立一个统一的图,融合基于视觉特征的图像相似性图和图像-标签二分图; 2)然后提出一种新的随机游走模型,该模型利用融合参数来平衡图像内容和标签之间的影响。此外,该框架不仅可以自然地结合伪相关反馈过程,但它也可以直接应用到应用程序,如基于内容的图像检索,基于文本的图像检索,和图像注释。在一个大的Flickr数据集上的实验分析表明,我们提出的框架的有效性和效率。
With the exponential growth of Web 2.0 applications, tags have been used extensively to describe the image contents on the Web. Due to the noisy and sparse nature in the human generated tags, how to understand and utilize these tags for image retrieval tasks has become an emerging research direction. As the low-level visual features can provide fruitful information, they are employed to improve the image retrieval results. However, it is challenging to bridge the semantic gap between image contents and tags. To attack this critical problem, we propose a unified framework in this paper which stems from a two-level data fusions between the image contents and tags: 1) A unified graph is built to fuse the visual feature-based image similarity graph with the image-tag bipartite graph; 2) A novel random walk model is then proposed, which utilizes a fusion parameter to balance the influences between the image contents and tags. Furthermore, the presented framework not only can naturally incorporate the pseudo relevance feedback process, but also it can be directly applied to applications such as content-based image retrieval, text-based image retrieval, and image annotation. Experimental analysis on a large Flickr dataset shows the effectiveness and efficiency of our proposed framework.