Supervised Distributed Hashing for Large-Scale Multimedia Retrieval

Supervised Distributed Hashing for Large-Scale Multimedia Retrieval
复制标题

DOI:
10.1109/tmm.2017.2749160
复制
发表时间:
2018-03
影响因子:
7.3
通讯作者:
Deming Zhai;Xianming Liu;Xiangyang Ji;Debin Zhao;S. Satoh;Wen Gao
Deming Zhai;Xianming Liu;Xiangyang Ji;Debin Zhao;S. Satoh;Wen Gao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Deming Zhai;Xianming Liu;Xiangyang Ji;Debin Zhao;S. Satoh;Wen Gao

文献摘要

被引文献

相似文献

近年来,哈希技术在大规模多媒体检索中的应用日益广泛。已经为存储在单个机器中的数据设计了广泛的散列方法,即集中式散列。然而,在许多现实世界的应用程序中,大规模数据通常分布在不同的位置、服务器或站点。虽然分布式数据的哈希在理论上可以通过将所有分布式数据组合成一个完整的数据集来实现,但在实践中通常会导致高昂的计算、通信和存储成本。到目前为止,只有几种方法是为分布式散列量身定做的,都是无监督的方法。本文提出了一种高效的有监督分布式散列方法(SupDisH),该方法通过分布式利用语义标签信息来学习区分散列函数。具体地说,我们将分布式哈希问题引入到分类框架中,期望学习的二进制代码足够不同,以便进行语义检索。通过引入辅助变量,将分布式模型分解为一组具有一致性约束的分散子问题,这些子问题可以在分布式网络的每个顶点上并行求解。这样,我们可以获得高质量的独特的无偏二进制码和低计算复杂度的一致哈希函数,这有助于处理涉及分布式数据集的大规模多媒体检索任务。在三个大规模数据集上的实验评估表明,SupDisH与集中式散列方法具有竞争优势,并且显著优于最新的无监督分布式方法。
Recent years have witnessed the growing popularity of hashing for large-scale multimedia retrieval. Extensive hashing methods have been designed for data stored in a single machine, that is, centralized hashing . In many real-world applications, however, the large-scale data are often distributed across different locations, servers, or sites. Although hashing for distributed data can be implemented by assembling all distributed data together as a whole dataset in theory, it usually leads to prohibitive computation, communication, and storage costs in practice. Up to now, only a few methods were tailored for distributed hashing, which are all unsupervised approaches. In this paper, we propose an efficient and effective method called supervised distributed hashing (SupDisH), which learns discriminative hash functions by leveraging the semantic label information in a distributed manner. Specifically, we cast the distributed hashing problem into the framework of classification, where the learned binary codes are expected to be distinct enough for semantic retrieval. By introducing auxiliary variables, the distributed model is then separated into a set of decentralized subproblems with consistency constraints, which can be solved in parallel on each vertex of the distributed network. As such, we can obtain high-quality distinctive unbiased binary codes and consistent hash functions with low computational complexity, which facilitate tackling large-scale multimedia retrieval tasks involving distributed datasets. Experimental evaluations on three large-scale datasets show that SupDisH is competitive to centralized hashing methods and outperforms the state-of-the-art unsupervised distributed method significantly.