Hybrid representation learning for cross-modal retrieval

Hybrid representation learning for cross-modal retrieval
复制标题

用于跨模态检索的混合表示学习

DOI:
10.1016/j.neucom.2018.10.082
复制
发表时间:
2019-06
期刊:
影响因子:
6
通讯作者:
He Zhiquan
He Zhiquan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Cao Wenming;Lin Qiubin;He Zhihai;He Zhiquan

文献摘要

参考文献

被引文献

相似文献

深度神经网络在单模态检索中的快速发展,推动了深度神经网络在跨模态检索任务中的广泛应用。因此,我们提出了一种基于DNN的方法来学习每个模态的共享表示。我们的方法,混合表示学习(HRL),包括三个步骤。在第一个学习步骤中,堆叠的限制玻尔兹曼机(SRBM)被用来提取模态友好的表示为每一个模态,与统计特性,这是更相似的比那些原始的输入实例的两个模态,和一个多模态深度信念网(多模态DBN)被用来提取模态相互表示,其中包含一些丢失的信息,在原始的输入实例。在第二步学习中,使用了一个包含联合自动编码器和一个三层前馈神经网络的两级网络。从这些步骤中,获得混合表示,其组合了由图像路径SRBM和模态互表示构建的图像表示,模态互表示涉及潜像表示,并且可以用于经由多模态DBN来推断图像的缺失值,反之亦然。在第三个学习步骤中,使用堆叠的双峰自动编码器来获得每个模态的最终共享表示。实验结果表明,我们提出的HRL方法是上级优于几个先进的方法,根据三个广泛使用的跨模态数据集。
The rapid development of Deep Neural Networks (DNNs) in single-modal retrieval has promoted the wide application of DNNs in cross-modal retrieval tasks. Therefore, we propose a DNN-based method to learn the shared representation for each modality. Our method, hybrid representation learning (HRL), consists of three steps. In the first learning step, stacked restricted Boltzmann machines (SRBM) are utilized to extract the modality-friendly representation for each modality, with statistical properties that are more similar than those of the original input instances of both modalities, and a multimodal deep belief net (multimodal DBN) is utilized to extract the modality-mutual representation, which contains some missing information in the original input instances. In the second learning step, a two-level network containing a joint autoencoder and a three-layer feedforward neural net are used. From these steps, the hybrid representation is obtained, which combines the image representation constructed by the image-pathway SRBM and the modality-mutual representation, which involves the latent image representation and can be used to infer the missing values of the image via the multimodal DBN or vice-versa. In the third learning step, stacked bimodal autoencoders are used to obtain the final shared representation for each modality. The experimental results show that our proposed HRL method is superior to several advanced approaches according to three widely used cross-modal datasets.
DOI: 10.1007/s11263-013-0658-4
发表时间: 2014-01-01
影响因子: 19.5
作者:
Gong, Yunchao;Ke, Qifa;Lazebnik, Svetlana
通讯作者: Lazebnik, Svetlana
DOI: 10.1016/j.neucom.2017.09.012
发表时间: 2018-01
期刊: Neurocomputing
影响因子: 6
作者:
Wei Zhang;Qi Chen;W. Zhang;Xuanyu He
通讯作者: Wei Zhang;Qi Chen;W. Zhang;Xuanyu He
使用稀疏和半监督正则化学习跨媒体联合表示
DOI: 10.1109/tcsvt.2013.2276704
发表时间: 2014-06-01
影响因子: 8.4
作者:
Zhai, Xiaohua;Peng, Yuxin;Xiao, Jianguo
通讯作者: Xiao, Jianguo
DOI: --
发表时间: 2016-12
期刊: ArXiv
影响因子: --
作者:
Marcel Simon;E. Rodner;Joachim Denzler
通讯作者: Marcel Simon;E. Rodner;Joachim Denzler
DOI: 10.1007/978-1-4612-4380-9_14
发表时间: 1936-12
期刊: Biometrika
影响因子: 2.7
作者:
H. Hotelling
通讯作者: H. Hotelling