How should we evaluate supervised hashing?

How should we evaluate supervised hashing?
复制标题

DOI:
10.1109/icassp.2017.7952453
复制
发表时间:
2016-09
期刊:
2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Alexandre Sablayrolles;Matthijs Douze;Nicolas Usunier;H. Jégou
Alexandre Sablayrolles;Matthijs Douze;Nicolas Usunier;H. Jégou
中科院分区:
其他
文献类型:
--
作者:
Alexandre Sablayrolles;Matthijs Douze;Nicolas Usunier;H. Jégou

文献摘要

被引文献

相似文献

散列生成文档的紧凑表示,以基于这些短代码执行分类或检索等任务。当散列被监督时,代码使用训练数据上的标签进行训练。本文首先表明,在文献中使用的评估协议监督哈希是不令人满意的:我们表明,一个平凡的解决方案,编码的分类器的输出显着优于现有的监督或半监督的方法,同时使用更短的代码。然后,我们提出了两种替代协议的监督哈希:一个基于检索的一组不相交的类,另一个基于转移学习到新的类。我们为图像相关任务提供了两种基线方法来评估(半)监督哈希的性能:无编码和无监督代码。这些基线给出了监督哈希方案性能的下限和上限。
Hashing produces compact representations for documents, to perform tasks like classification or retrieval based on these short codes. When hashing is supervised, the codes are trained using labels on the training data. This paper first shows that the evaluation protocols used in the literature for supervised hashing are not satisfactory: we show that a trivial solution that encodes the output of a classifier significantly outperforms existing supervised or semi-supervised methods, while using much shorter codes. We then propose two alternative protocols for supervised hashing: one based on retrieval on a disjoint set of classes, and another based on transfer learning to new classes. We provide two baseline methods for image-related tasks to assess the performance of (semi-)supervised hashing: without coding and with unsupervised codes. These baselines give a lower- and upper-bound on the performance of a supervised hashing scheme.