A Multi-View Embedding Space for Modeling Internet Images, Tags, and Their Semantics

A Multi-View Embedding Space for Modeling Internet Images, Tags, and Their Semantics
复制标题

DOI:
10.1007/s11263-013-0658-4
复制
发表时间:
2014-01-01
影响因子:
19.5
通讯作者:
Lazebnik, Svetlana
Lazebnik, Svetlana
中科院分区:
计算机科学2区
文献类型:
--
作者:
Gong, Yunchao;Ke, Qifa;Lazebnik, Svetlana

文献摘要

被引文献

相似文献

本文研究了用于图像到图像搜索、标签到图像搜索和图像到标签搜索(图像注释)等任务的互联网图像和相关文本或标签的建模问题。我们从典型相关分析(CCA)开始,这是一种流行且成功的方法,用于将视觉和文本特征映射到相同的潜在空间,并加入了第三个视图,该视图捕捉由单个类别或多个非互斥概念表示的高级图像语义。我们提出了两种训练三视图嵌入的方法:有监督的,第三视图来自地面事实标签或搜索关键字;以及非监督的,通过聚类标签自动获得语义主题。为了在保持学习过程可伸缩性的同时保证检索任务的高准确率,我们结合了多个强视觉特征,并使用显式非线性核映射来有效地逼近核CCA。为了进行检索,我们在嵌入空间中使用了一个专门设计的相似性函数,该函数的性能大大优于欧几里德距离。由此产生的系统产生了令人信服的定性结果,并在三个大规模互联网图像数据集上的检索任务上超过了一些两视图基线。
This paper investigates the problem of modeling Internet images and associated text or tags for tasks such as image-to-image search, tag-to-image search, and image-to-tag search (image annotation). We start with canonical correlation analysis (CCA), a popular and successful approach for mapping visual and textual features to the same latent space, and incorporate a third view capturing high-level image semantics, represented either by a single category or multiple non-mutually-exclusive concepts. We present two ways to train the three-view embedding: supervised, with the third view coming from ground-truth labels or search keywords; and unsupervised, with semantic themes automatically obtained by clustering the tags. To ensure high accuracy for retrieval tasks while keeping the learning process scalable, we combine multiple strong visual features and use explicit nonlinear kernel mappings to efficiently approximate kernel CCA. To perform retrieval, we use a specially designed similarity function in the embedded space, which substantially outperforms the Euclidean distance. The resulting system produces compelling qualitative results and outperforms a number of two-view baselines on retrieval tasks on three large-scale Internet image datasets.