Composite hashing with multiple information sources

Composite hashing with multiple information sources
复制标题

DOI:
10.1145/2009916.2009950
复制
发表时间:
2011-07
期刊:
Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval
影响因子:
--
通讯作者:
Dan Zhang;Fei Wang;Luo Si
Dan Zhang;Fei Wang;Luo Si
中科院分区:
其他
文献类型:
--
作者:
Dan Zhang;Fei Wang;Luo Si

文献摘要

被引文献

相似文献

具有大量文本和图像数据的相似性搜索应用需要一种高效的解决方案。一种有用的策略是通过语义哈希将数据库中的示例表示为紧凑的二进制代码,由于其快速的查询/搜索速度和大幅降低的存储需求而引起了人们的广泛关注。目前所有的语义哈希方法都只处理每个示例由一种类型的特征表示的情况。然而,在许多真实的应用程序中,通常从几个不同的信息源描述示例。例如,网页的特征可以从其内容部分及其相关联的链接两者导出。为了解决在这种情况下学习好的哈希码的问题,我们提出了一个新的研究问题--多信息源复合哈希(CHMIS)。新的研究问题的焦点是设计一种算法,将来自不同信息源的特征有效地合并到二进制哈希码中。特别是,我们提出了一个算法CHMIS-AW(CHMIS与调整权重)学习的代码。该算法通过调整每个源的权重将多个不同源的信息集成到二进制哈希码中,以最大限度地提高编码性能,并实现从查询示例到其二进制哈希码的快速转换。在五个不同数据集上的实验结果表明,该方法与其他几种最先进的语义哈希技术相比具有上级性能。
Similarity search applications with a large amount of text and image data demands an efficient and effective solution. One useful strategy is to represent the examples in databases as compact binary codes through semantic hashing, which has attracted much attention due to its fast query/search speed and drastically reduced storage requirement. All of the current semantic hashing methods only deal with the case when each example is represented by one type of features. However, examples are often described from several different information sources in many real world applications. For example, the characteristics of a webpage can be derived from both its content part and its associated links. To address the problem of learning good hashing codes in this scenario, we propose a novel research problem -- Composite Hashing with Multiple Information Sources (CHMIS). The focus of the new research problem is to design an algorithm for incorporating the features from different information sources into the binary hashing codes efficiently and effectively. In particular, we propose an algorithm CHMIS-AW (CHMIS with Adjusted Weights) for learning the codes. The proposed algorithm integrates information from several different sources into the binary hashing codes by adjusting the weights on each individual source for maximizing the coding performance, and enables fast conversion from query examples to their binary hashing codes. Experimental results on five different datasets demonstrate the superior performance of the proposed method against several other state-of-the-art semantic hashing techniques.