Canonical contextual distance for large-scale image annotation and retrieval

Canonical contextual distance for large-scale image annotation and retrieval
复制标题

DOI:
10.1145/1631058.1631062
复制
发表时间:
2009-10
期刊:
--
影响因子:
--
通讯作者:
Hideki Nakayama;T. Harada;Y. Kuniyoshi
Hideki Nakayama;T. Harada;Y. Kuniyoshi
中科院分区:
其他
文献类型:
--
作者:
Hideki Nakayama;T. Harada;Y. Kuniyoshi

文献摘要

相似文献

为了实现通用图像识别,系统需要学习世界上大量的目标及其外观。因此,利用大量的Web图像获取视觉知识的研究,最近,和基于搜索的方法,现在在这个研究领域蓬勃发展。然而,在一般情况下,这种方法的搜索过程中进行的相似性措施的基础上,简单的图像特征和遭受的语义差距。这是一个大问题,可能是整个系统的瓶颈。本文提出了一种基于新的相似性度量标准--典型上下文距离的图像标注和检索方法。该方法有效地利用了从多个标签中估计的图像上下文,并学习了基本的和有区别的潜在空间。使用概率结构,我们的相似性度量可以同时反映样本的外观和语义。由于我们的学习方法具有高度的可扩展性,因此它甚至在大型网络规模的数据集中也是有效的。因此,我们的相似性度量将有助于许多其他基于搜索的方法。在实验中,我们表明,我们的方法优于以前的作品使用标准的Corel基准。接下来,我们通过将其应用于350万张网络图像来验证我们的方法。
To realize generic image recognition, the system needs to learn an enormous amount of targets in the world and their appearances. Therefore, visual knowledge acquisition using massive amounts of web images has been studied recently, and search-based methods are now flourishing in this research field. However, in general, search process of such methods are conducted using similarity measures based on simple image features and suffer from the semantic-gap. This is a big problem and can be a bottleneck of the entire systems. In this paper, we propose a method of image annotation and retrieval based on the new similarity measure, Canonical Contextual Distance. This method effectively uses contexts of images estimated from multiple labels and learns the essential and discriminative latent space. Using the probabilistic structure, our similarity measure can reflect both appearance and semantics of samples. Because our learning method is highly scalable, it is even effective in a large web-scale dataset. Therefore, our similarity measure will be helpful to many other search-based methods. In the experiment, we show that our method outperforms previous works using the standard Corel benchmark. Next, we verify our method by applying it to 3.5 million web images.