Generalized Semantic Preserving Hashing for Cross-Modal Retrieval

Generalized Semantic Preserving Hashing for Cross-Modal Retrieval
复制标题

DOI:
10.1109/tip.2018.2863040
复制
发表时间:
2019-01
影响因子:
10.6
通讯作者:
Devraj Mandal;K. Chaudhury;S. Biswas
Devraj Mandal;K. Chaudhury;S. Biswas
中科院分区:
计算机科学1区
文献类型:
--
作者:
Devraj Mandal;K. Chaudhury;S. Biswas

文献摘要

被引文献

相似文献

由于大量多媒体数据的可用性,跨模态检索变得越来越重要。当数据量很大时,基于散列的技术为这个问题提供了一个有吸引力的解决方案。对于跨模态检索,来自两个模态的数据可以与单个标签或多个标签相关联,并且此外,可以具有或可以不具有一一对应关系。这项工作提出了一个简单的哈希框架,它有能力与不同的情况下,同时有效地捕捉数据项之间的语义关系。工作分两个阶段进行,其中第一阶段通过分解使用标签信息构造的亲和矩阵来学习最佳散列码。在第二阶段,岭回归和核逻辑回归用于学习用于将输入数据映射到比特域的散列函数。我们还提出了一种新的迭代解决方案的情况下,训练数据是非常大的,或者当整个训练数据是不可用的一次。在Wiki等单标签数据集和MirFlickr、NUS-WIDE、Pascal和LabelMe等多标签数据集上进行了大量实验,并与最先进的方法进行了比较,结果表明该方法是有效的。
Cross-modal retrieval is gaining importance due to the availability of large amounts of multimedia data. Hashing-based techniques provide an attractive solution to this problem when the data size is large. For cross-modal retrieval, data from the two modalities may be associated with a single label or multiple labels, and in addition, may or may not have a one-to-one correspondence. This work proposes a simple hashing framework which has the capability to work with different scenarios while effectively capturing the semantic relationship between the data items. The work proceeds in two stages in which the first stage learns the optimum hash codes by factorizing an affinity matrix, constructed using the label information. In the second stage, ridge regression and kernel logistic regression is used to learn the hash functions for mapping the input data to the bit domain. We also propose a novel iterative solution for cases where the training data is very large, or when the whole training data is not available at once. Extensive experiments on single label data set like Wiki and multi-label datasets like MirFlickr, NUS-WIDE, Pascal, and LabelMe, and comparisons with the state-of-the-art, shows the usefulness of the proposed approach.