Sequential Discrete Hashing for Scalable Cross-Modality Similarity Retrieval

Sequential Discrete Hashing for Scalable Cross-Modality Similarity Retrieval
复制标题

DOI:
10.1109/tip.2016.2619262
复制
发表时间:
2017
影响因子:
10.6
通讯作者:
Li Liu;Zijia Lin;Ling Shao;Fumin Shen;Guiguang Ding;J. Han
Li Liu;Zijia Lin;Ling Shao;Fumin Shen;Guiguang Ding;J. Han
中科院分区:
计算机科学1区
文献类型:
--
作者:
Li Liu;Zijia Lin;Ling Shao;Fumin Shen;Guiguang Ding;J. Han

文献摘要

被引文献

相似文献

随着互联网的飞速发展,如何利用大规模的多模式Web数据检索技术已经成为计算机视觉和多媒体领域最热门但也是最具挑战性的问题之一。最近,散列方法被用于大规模数据空间中的快速最近邻搜索,方法是将高维特征描述符嵌入到低维的保持相似性的Hamming空间中。受此启发,本文提出了一种新的有监督的跨通道散列框架,该框架可以为不同通道表示的实例生成统一的二进制代码。具体地,在学习阶段,可以利用基于提升策略联合最小化其经验损失的离散优化方案来顺序地学习代码的每一位。然后,以按位方式学习每个通道的散列函数,将对应的表示映射到统一的散列码。我们将这种方法称为跨通道顺序离散哈希(CSDH),它可以有效地减少过度简化的舍入步骤中产生的量化误差,从而产生高质量的二进制代码。在测试阶段,通过合并来自不同模态的未见实例的预测散列结果,利用简单的融合方案来生成用于最终检索的统一散列码。在Wiki、MIRFlickr和NUS-Wide三个标准数据集上对所提出的CSDH进行了系统的评估,结果表明,我们的方法的性能明显优于最先进的多模式哈希技术。
With the dramatic development of the Internet, how to exploit large-scale retrieval techniques for multimodal web data has become one of the most popular but challenging problems in computer vision and multimedia. Recently, hashing methods are used for fast nearest neighbor search in large-scale data spaces, by embedding high-dimensional feature descriptors into a similarity preserving Hamming space with a low dimension. Inspired by this, in this paper, we introduce a novel supervised cross-modality hashing framework, which can generate unified binary codes for instances represented in different modalities. Particularly, in the learning phase, each bit of a code can be sequentially learned with a discrete optimization scheme that jointly minimizes its empirical loss based on a boosting strategy. In a bitwise manner, hash functions are then learned for each modality, mapping the corresponding representations into unified hash codes. We regard this approach as cross-modality sequential discrete hashing (CSDH), which can effectively reduce the quantization errors arisen in the oversimplified rounding-off step and thus lead to high-quality binary codes. In the test phase, a simple fusion scheme is utilized to generate a unified hash code for final retrieval by merging the predicted hashing results of an unseen instance from different modalities. The proposed CSDH has been systematically evaluated on three standard data sets: Wiki, MIRFlickr, and NUS-WIDE, and the results show that our method significantly outperforms the state-of-the-art multimodality hashing techniques.