Deep robust multilevel semantic hashing for multi-label cross-modal retrieval

Deep robust multilevel semantic hashing for multi-label cross-modal retrieval
复制标题

DOI:
10.1016/j.patcog.2021.108084
复制
发表时间:
2021-06
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Ge Song;Xiaoyang Tan;Jun Zhao;Ming Yang
Ge Song;Xiaoyang Tan;Jun Zhao;Ming Yang
中科院分区:
其他
文献类型:
--
作者:
Ge Song;Xiaoyang Tan;Jun Zhao;Ming Yang

文献摘要

相似文献

基于散列的跨模式检索最近取得了重大进展。但是,由于固有的通道差异和噪声,直接将包含丰富语义的不同通道的数据嵌入到联合汉明空间中,不可避免地会产生伪码。为提高多标签跨模式检索的精确度,提出了一种深度稳健的多级语义哈希算法。它寻求保持语义丰富的数据之间的细粒度相似性,即多标签,而显式地要求相异点之间的距离大于特定值以获得较强的健壮性。为此,我们在信息编码理论分析的基础上给出了该值的一个有效界,并将上述目标体现为一种差值自适应的三元组损失。此外,我们通过融合多个哈希码引入伪码来挖掘稀有语义,缓解了相似信息的稀疏性问题。在三个基准测试上的实验表明了所得界的有效性,并且我们的方法达到了最先进的性能。
Hashing based cross-modal retrieval has recently made significant progress. But straightforward embedding data from different modalities involving rich semantics into a joint Hamming space will inevitably produce false codes due to the intrinsic modality discrepancy and noises. We present a novel deep Robust Multilevel Semantic Hashing (RMSH) for more accurate multi-label cross-modal retrieval. It seeks to preserve fine-grained similarity among data with rich semantics,i.e., multi-label, while explicitly require distances between dissimilar points to be larger than a specific value for strong robustness. For this, we give an effective bound of this value based on the information coding-theoretic analysis, and the above goals are embodied into a margin-adaptive triplet loss. Furthermore, we introduce pseudo-codes via fusing multiple hash codes to explore seldom-seen semantics, alleviating the sparsity problem of similarity information. Experiments on three benchmarks show the validity of the derived bounds, and our method achieves state-of-the-art performance.