Deep Multimodal Complementarity Learning

Deep Multimodal Complementarity Learning
复制标题

DOI:
10.1109/tnnls.2022.3165180
复制
发表时间:
2022-04
影响因子:
10.4
通讯作者:
Daheng Wang;Tong Zhao;Wenhao Yu;N. Chawla;Meng Jiang
Daheng Wang;Tong Zhao;Wenhao Yu;N. Chawla;Meng Jiang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Daheng Wang;Tong Zhao;Wenhao Yu;N. Chawla;Meng Jiang

文献摘要

相似文献

互补性在复杂数据对象的不同组件所产生的协同效应中起着重要作用。多模态数据的互补性学习是表征学习的基本挑战,因为互补性存在于多个模态和每个模态的一个或多个项目中。此外,还需要一个合适的度量来度量表示空间中的互补性。依赖于基于相似性的度量的现有方法不能充分捕获互补性。在这项工作中,我们提出了一种新的深度架构,用于系统地从多模态多项目数据中学习组件的互补性。该模型包括三个主要模块:1)单峰聚合,用于提取模内互补性;2)跨模态融合,提取模态层面的多模态互补性;3)交互聚合,在项目层面提取多式联运互补性。为了量化互补性,我们利用TUBE距离度量来度量合成数据对象与其在表示空间中的标签之间的差异。在三个真实数据集上的实验表明,我们的模型在对象分类上的平均倒数秩(MRR)优于最先进的+6.8%,在保留项目预测上的MRR优于最先进的+3.0%。定性分析表明,互补性与相似性存在显著差异。
Complementarity plays a significant role in the synergistic effect created by different components of a complex data object. Complementarity learning on multimodal data has fundamental challenges of representation learning because the complementarity exists along with multiple modalities and one or multiple items of each modality. Also, an appropriate metric is needed for measuring the complementarity in the representation space. Existing methods that rely on similarity-based metrics cannot adequately capture the complementarity. In this work, we propose a novel deep architecture for systematically learning the complementarity of components from multimodal multi-item data. The proposed model consists of three major modules: 1) unimodal aggregation for extracting the intramodal complementarity; 2) cross-modal fusion for extracting the intermodal complementarity at the modality level; and 3) interactive aggregation for extracting the intermodal complementarity at the item level. To quantify complementarity, we utilize the TUBE distance metric to measure the difference between the composited data object and its label in the representation space. Experiments on three real datasets show that our model outperforms the state-of-the-art by +6.8% of mean reciprocal rank (MRR) on object classification and +3.0% of MRR on hold-out item prediction. Qualitative analyses reveal that complementarity is significantly different from similarity.