Mining on Heterogeneous Manifolds for Zero-Shot Cross-Modal Image Retrieval

Mining on Heterogeneous Manifolds for Zero-Shot Cross-Modal Image Retrieval
复制标题

DOI:
10.1609/aaai.v34i07.6949
复制
发表时间:
2020-04
期刊:
--
影响因子:
--
通讯作者:
Fan Yang;Zheng Wang-;Jing Xiao;S. Satoh
Fan Yang;Zheng Wang-;Jing Xiao;S. Satoh
中科院分区:
其他
文献类型:
--
作者:
Fan Yang;Zheng Wang-;Jing Xiao;S. Satoh

文献摘要

被引文献

相似文献

最新的零拍摄跨模态图像检索方法将来自不同模态的图像映射到统一的特征空间中,通过使用预训练的模型来利用它们的相关性。基于零拍摄图像的流形通常是变形和不完整的观察,我们认为,在双流模型的训练过程中,看不见的类的流形不可避免地会被扭曲,双流模型只是将来自不同模态的图像映射到一个统一的空间。这个问题直接导致跨模态检索性能差。我们提出了一种双向随机游走方案,通过遍历每个模态的特征空间中的异质流形来挖掘图像之间更可靠的关系。我们所提出的方法受益于模态内分布,以减轻由噪声相似性在跨模态特征空间中造成的干扰。因此,我们在热性能方面取得了很大的改进,可见图像检索任务。本文代码:https://github.com/fyang93/cross-modal-retrieval
Most recent approaches for the zero-shot cross-modal image retrieval map images from different modalities into a uniform feature space to exploit their relevance by using a pre-trained model. Based on the observation that manifolds of zero-shot images are usually deformed and incomplete, we argue that the manifolds of unseen classes are inevitably distorted during the training of a two-stream model that simply maps images from different modalities into a uniform space. This issue directly leads to poor cross-modal retrieval performance. We propose a bi-directional random walk scheme to mining more reliable relationships between images by traversing heterogeneous manifolds in the feature space of each modality. Our proposed method benefits from intra-modal distributions to alleviate the interference caused by noisy similarities in the cross-modal feature space. As a result, we achieved great improvement in the performance of the thermal v.s. visible image retrieval task. The code of this paper: https://github.com/fyang93/cross-modal-retrieval