Cross-modal dual subspace learning with adversarial network

Cross-modal dual subspace learning with adversarial network
复制标题

DOI:
10.1016/j.neunet.2020.03.015
复制
发表时间:
2020-03
期刊:
Neural networks : the official journal of the International Neural Network Society
影响因子:
--
通讯作者:
Fei Shang;Huaxiang Zhang;Jiande Sun;Liqiang Nie;Li Liu
Fei Shang;Huaxiang Zhang;Jiande Sun;Liqiang Nie;Li Liu
中科院分区:
其他
文献类型:
--
作者:
Fei Shang;Huaxiang Zhang;Jiande Sun;Liqiang Nie;Li Liu

文献摘要

相似文献

近年来,随着多模式数据的快速发展,跨模式检索引起了人们的极大兴趣,有效利用不同模式数据之间的互补关系和尽可能消除异构性鸿沟是跨模式检索面临的两个关键挑战。提出了一种基于对抗性网络的跨模式对偶子空间学习(DSAN)网络模型。本文的主要贡献如下:(1)提出了双重子空间(视觉子空间和文本子空间),能够更好地挖掘不同通道的底层结构信息和通道特有信息。(2)考虑了正负样本之间的相对距离和绝对距离,并引入了硬样本挖掘的思想,提出了一种改进的四元组损失。(3)为了最大化最相似的跨模负样本与其对应的跨模正样本之间的距离,提出了模内约束损失的概念。特别是,特征保留和通道分类就像是两个对立面。DSAN试图缩小不同模式之间的异质差距,并在对偶子空间中区分随机样本的原始模式。综合实验结果表明,DSAN在4个跨通道数据集上的性能明显优于9种最先进的方法。
Cross-modal retrieval has recently attracted much interest along with the rapid development of multimodal data, and effectively utilizing the complementary relationship of different modal data and eliminating the heterogeneous gap as much as possible are the two key challenges. In this paper, we present a novel network model termed cross-modal Dual Subspace learning with Adversarial Network (DSAN). The main contributions are as follows: (1) Dual subspaces (visual subspace and textual subspace) are proposed, which can better mine the underlying structure information of different modalities as well as modality-specific information. (2) An improved quadruplet loss is proposed, which takes into account the relative distance and absolute distance between positive and negative samples, together with the introduction of the idea of hard sample mining. (3) Intra-modal constrained loss is proposed to maximize the distance of the most similar cross-modal negative samples and their corresponding cross-modal positive samples. In particular, feature preserving and modality classification act as two antagonists. DSAN tries to narrow the heterogeneous gap between different modalities, and distinguish the original modality of random samples in dual subspaces. Comprehensive experimental results demonstrate that, DSAN significantly outperforms 9 state-of-the-art methods on four cross-modal datasets.