Deep continual hashing with gradient-aware memory for cross-modal retrieval

Deep continual hashing with gradient-aware memory for cross-modal retrieval
复制标题

DOI:
10.1016/j.patcog.2022.109276
复制
发表时间:
2022-12
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Ge Song;Xiaoyang Tan;Ming Yang
Ge Song;Xiaoyang Tan;Ming Yang
中科院分区:
其他
文献类型:
--
作者:
Ge Song;Xiaoyang Tan;Ming Yang

文献摘要

相似文献

跨模态哈希算法(CMH)已广泛应用于大规模多媒体检索。然而,大多数当前CMH方法关注于封闭的检索场景,而不是真实世界的环境,即,复杂和不断变化的语义。当包含新类对象的数据出现时,当前的CMH必须在所有历史训练数据上重新训练模型,而不是新数据,以适应新的语义,但互联网上永不停止的数据上传使得这不切实际。在本文中,我们设计了一种深度哈希方法,称为连续跨模态哈希与梯度感知存储器(CCMH-GAM),用于学习类别增加的多标签跨模态数据的二进制代码。CCMH-GAM是一个两步哈希架构,一个哈希网络学习哈希数据的增加语义,即,标签映射到语义代码中,而其他特定于模态的散列网络学习将数据映射到相应的语义代码中。具体地说,为了保持旧语义的编码能力,对前一个网络设计了一种基于累积低存储标签-代码对的正则化。对于特定模态网络,我们提出了一种记忆构造方法,通过一些样本来近似所有数据的全情节梯度,并推导出其快速实现与近似误差的上限。基于这种记忆,我们提出了一种梯度投影方法,从理论上提高了模型更新后旧数据编码不变的概率。在三个数据集上的大量实验表明,CCMH-GAM可以不断学习哈希函数,并产生最先进的检索性能。
Cross-modal hashing (CMH) has become widely used for large-scale multimedia retrieval. However, most current CMH methods focus on the closed retrieval scenario, not the real-world environments, i.e., complex and changing semantics. When data containing new class objects emerge, the current CMH has to retrain the model on all history training data, not the new data, to accommodate new semantics, but the never-stop upload of data on the Internet makes this impractical. In this paper, we devise a deep hashing method called Continual Cross-Modal Hashing with Gradient Aware Memory (CCMH-GAM) for learning binary codes of multi-label cross-modal data with increasing categories. CCMH-GAM is a two-step hashing architecture, one hashing network learns to hash the increasing semantics of data, i.e., label, into the semantic codes, and other modality-specific hashing networks learn to map data into the corresponding semantic codes. Specifically, to keep the encoding ability for old semantics, a regularization based on accumulating low-storage label-code pairs is designed for the former network. For the modality-specific networks, we propose a memory construction method via approximating the full episodic gradients of all data by some exemplars and derive its fast implementation with the upper bound of approximation error. Based on this memory, we propose a gradient projection method to theoretically improve the probability of old data’s code being unchanged after updating the model. Extensive experiments on three datasets demonstrate that CCMH-GAM can continually learn hash functions and yield state-of-the-art retrieval performance.