Self-Supervised Video Hashing With Hierarchical Binary Auto-Encoder

Self-Supervised Video Hashing With Hierarchical Binary Auto-Encoder
复制标题

使用分层二进制自动编码器进行自监督视频哈希

DOI:
10.1109/tip.2018.2814344
复制
发表时间:
2018-07-01
影响因子:
10.6
通讯作者:
Hong, Richang
Hong, Richang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Song, Jingkuan;Zhang, Hanwang;Hong, Richang

文献摘要

被引文献

相似文献

现有的视频哈希函数建立在三个独立的阶段:帧汇集、松弛学习和二值化,在联合二进制优化模型中没有充分探索视频帧的时间顺序,导致严重的信息损失。本文提出了一种新的无监督视频哈希框架,称为自监督视频哈希(SSVH),它能够以端到端的学习哈希方式捕获视频的时间特性。我们具体解决了两个核心问题:1)如何设计一个编解码器结构来生成视频的二进制代码;2)如何使二进制代码具有准确的视频检索能力。我们设计了一种分层的二进制自动编码器来对视频中的时间依赖关系进行建模,并将视频嵌入到二进制码中,比堆叠结构的计算量更少。然后,我们鼓励二进制编码同时重建视频的视觉内容和邻域结构。在两个真实数据集上的实验表明,我们的SSVH方法在非监督视频检索任务上的性能明显优于现有的方法,并且取得了目前最好的性能。
Existing video hash functions are built on three isolated stages: frame pooling, relaxed learning, and binarization, which have not adequately explored the temporal order of video frames in a joint binary optimization model, resulting in severe information loss. In this paper, we propose a novel unsupervised video hashing framework dubbed self-supervised video hashing (SSVH), which is able to capture the temporal nature of videos in an end-to-end learning to hash fashion. We specifically address two central problems: 1) how to design an encoder–decoder architecture to generate binary codes for videos and 2) how to equip the binary codes with the ability of accurate video retrieval. We design a hierarchical binary auto-encoder to model the temporal dependencies in videos with multiple granularities, and embed the videos into binary codes with less computations than the stacked architecture. Then, we encourage the binary codes to simultaneously reconstruct the visual content and neighborhood structure of the videos. Experiments on two real-world data sets show that our SSVH method can significantly outperform the state-of-the-art methods and achieve the current best performance on the task of unsupervised video retrieval.