Hierarchical Multimodal Fusion Network with Dynamic Multi-task Learning

Hierarchical Multimodal Fusion Network with Dynamic Multi-task Learning
复制标题

DOI:
10.1109/iri51335.2021.00034
复制
发表时间:
2021-08
期刊:
2021 IEEE 22nd International Conference on Information Reuse and Integration for Data Science (IRI)
影响因子:
--
通讯作者:
Tianyi Wang;Shu‐Ching Chen
Tianyi Wang;Shu‐Ching Chen
中科院分区:
其他
文献类型:
--
作者:
Tianyi Wang;Shu‐Ching Chen

文献摘要

相似文献

真实世界的数据通常包含多个模态和非排他性标签。多模态融合是多模态学习中的一个重要步骤,它将来自不同模态的特征整合到向量空间中,使得分类器可以利用融合后的向量来生成最终的预测分数。常见的多模态融合方法很少考虑跨模态交互,而跨模态交互在利用模态间关系和随后创建联合模态嵌入方面起着至关重要的作用。在本文中,我们提出了一个分层多模态融合框架与动态多任务学习。它专注于为所有跨模态交互建模联合嵌入空间,并调整任务损失以获得最佳性能。该模型使用了一种新的分层多模态融合网络,该网络学习所有模态组合之间的跨模态交互,并以样本感知的方式动态分配每对模态的权重。此外,一种新的动态多任务学习方法被应用到处理多标签问题,通过自动调整任务级和样本级的学习进度。我们表明,所提出的框架优于基线和一些国家的最先进的方法。我们还展示了所提出的分层多模态融合和动态多任务学习单元的灵活性和模块化,可应用于各种类型的网络。
Real-world data often contain multiple modalities and non-exclusive labels. Multimodal fusion is a vital step in mul-timodallearning that integrates features from various modalities in the vector space so that the classifier could utilize the fused vector to generate the final prediction score. Common multimodal fusion approaches rarely consider the cross-modality interactions which play an essential role in exploiting the inter-modality relationship and subsequently creating the joint modality embedding. In this paper, we propose a hierarchical multimodal fusion framework with dynamic multi-task learning. It focuses on modeling the joint embedding space for all cross-modality interactions and adjusting the task loss for optimal performance. The proposed model uses a novel hierarchical multimodal fusion network that learns the cross-modal interactions among all combinations of modalities and dynamically allocates the weights for each pair in a sample-aware fashion. Furthermore, a novel dynamic multi-task learning approach is applied to handle the multi-label problems by automatically adjusting the learning progress on both task level and sample level. We show that the proposed framework outperforms the baselines and some of the state-of-the-art methods. We also demonstrate the flexibility and modularity of the proposed hierarchical multimodal fusion and dynamic multi-task learning units, which can be applied to various types of networks.