Webly Supervised Joint Embedding for Cross-Modal Image-Text Retrieval

Webly Supervised Joint Embedding for Cross-Modal Image-Text Retrieval
复制标题

DOI:
10.1145/3240508.3240712
复制
发表时间:
2018-08
期刊:
Proceedings of the 26th ACM international conference on Multimedia
影响因子:
--
通讯作者:
Niluthpol Chowdhury Mithun;Rameswar Panda;E. Papalexakis;A. Roy-Chowdhury
Niluthpol Chowdhury Mithun;Rameswar Panda;E. Papalexakis;A. Roy-Chowdhury
中科院分区:
其他
文献类型:
--
作者:
Niluthpol Chowdhury Mithun;Rameswar Panda;E. Papalexakis;A. Roy-Chowdhury

文献摘要

被引文献

相似文献

视觉数据和自然语言描述之间的跨模式检索仍然是多媒体领域的一个长期挑战。虽然最近的图像-文本检索方法通过学习跨模式对齐的深层表示而提供了巨大的前景,但大多数这些方法都受到用覆盖有限数量的具有基本事实句子的图像的小规模数据集进行训练的问题的困扰。此外,通过用句子标注数百万张图像来创建更大的数据集是极其昂贵的,而且可能会导致模型有偏见。受深度神经网络中Weble监督学习最近的成功启发,我们利用随处可得的带有噪声标注的网络图像来学习稳健的图文联合表示。具体地说,我们的主要想法是在学习视觉-语义联合嵌入的培训中利用网络图像和相应的标签,以及完全注释的数据集。我们提出了一种两阶段的方法,该方法可以用弱标注的网络图像来增强典型的基于监督成对排序损失的公式,以学习更健壮的视觉语义嵌入。在两个标准基准数据集上的实验表明,与最先进的方法相比,我们的方法在图文检索方面取得了显著的性能提升。
Cross-modal retrieval between visual data and natural language description remains a long-standing challenge in multimedia. While recent image-text retrieval methods offer great promise by learning deep representations aligned across modalities, most of these methods are plagued by the issue of training with small-scale datasets covering a limited number of images with ground-truth sentences. Moreover, it is extremely expensive to create a larger dataset by annotating millions of images with sentences and may lead to a biased model. Inspired by the recent success of webly supervised learning in deep neural networks, we capitalize on readily-available web images with noisy annotations to learn robust image-text joint representation. Specifically, our main idea is to leverage web images and corresponding tags, along with fully annotated datasets, in training for learning the visual-semantic joint embedding. We propose a two-stage approach for the task that can augment a typical supervised pair-wise ranking loss based formulation with weakly-annotated web images to learn a more robust visual-semantic embedding. Experiments on two standard benchmark datasets demonstrate that our method achieves a significant performance gain in image-text retrieval compared to state-of-the-art approaches.