Latent semantic sparse hashing for cross-modal similarity search

Latent semantic sparse hashing for cross-modal similarity search
复制标题

DOI:
10.1145/2600428.2609610
复制
发表时间:
2014-07
期刊:
Proceedings of the 37th international ACM SIGIR conference on Research & development in information retrieval
影响因子:
--
通讯作者:
J. Zhou;Guiguang Ding;Yuchen Guo
J. Zhou;Guiguang Ding;Yuchen Guo
中科院分区:
其他
文献类型:
--
作者:
J. Zhou;Guiguang Ding;Yuchen Guo

文献摘要

被引文献

相似文献

基于散列的相似性搜索方法能够在包含大量文本和图像的大规模多媒体数据库中进行高效的跨模态检索,已经引起了广泛的关注。跨模态散列算法的核心问题是如何在散列函数学习过程中有效地构建内在异构的多模态表示之间的相关性。与典型相关分析(CCA)类似,现有的跨模式哈希方法大多通过线性投影将异构数据嵌入到联合抽象空间中。然而,这些方法未能更有效地弥合语义鸿沟,并捕捉高层次的潜在语义信息,已被证明它可以导致更好的图像检索性能。为了解决这些挑战,在本文中,我们提出了一种新的潜在语义稀疏哈希(LSSH)执行跨模态相似性搜索,采用稀疏编码和矩阵分解。特别是,LSSH使用稀疏编码来捕获图像的显着结构,并使用矩阵分解来学习文本中的潜在概念。然后将学习到的潜在语义特征映射到一个联合抽象空间。此外,一个迭代策略被应用到有效地获得最优解,它有助于LSSH探索多模态表示之间的相关性,高效和自动。最后,通过高层抽象空间量化生成统一的哈希码。在三个不同的数据集上进行的大量实验突出了我们的方法在跨模态场景下的优势,并表明LSSH显着优于几种最先进的方法。
Similarity search methods based on hashing for effective and efficient cross-modal retrieval on large-scale multimedia databases with massive text and images have attracted considerable attention. The core problem of cross-modal hashing is how to effectively construct correlation between multi-modal representations which are heterogeneous intrinsically in the process of hash function learning. Analogous to Canonical Correlation Analysis (CCA), most existing cross-modal hash methods embed the heterogeneous data into a joint abstraction space by linear projections. However, these methods fail to bridge the semantic gap more effectively, and capture high-level latent semantic information which has been proved that it can lead to better performance for image retrieval. To address these challenges, in this paper, we propose a novel Latent Semantic Sparse Hashing (LSSH) to perform cross-modal similarity search by employing Sparse Coding and Matrix Factorization. In particular, LSSH uses Sparse Coding to capture the salient structures of images, and Matrix Factorization to learn the latent concepts from text. Then the learned latent semantic features are mapped to a joint abstraction space. Moreover, an iterative strategy is applied to derive optimal solutions efficiently, and it helps LSSH to explore the correlation between multi-modal representations efficiently and automatically. Finally, the unified hashcodes are generated through the high level abstraction space by quantization. Extensive experiments on three different datasets highlight the advantage of our method under cross-modal scenarios and show that LSSH significantly outperforms several state-of-the-art methods.