BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship Detection

BLOCK: Bilinear Superdiagonal Fusion for Visual Question Answering and Visual Relationship Detection
复制标题

DOI:
10.1609/aaai.v33i01.33018102
复制
发表时间:
2019-01
期刊:
ArXiv
影响因子:
--
通讯作者:
H. Ben-younes;Rémi Cadène;Nicolas Thome;M. Cord
H. Ben-younes;Rémi Cadène;Nicolas Thome;M. Cord
中科院分区:
其他
文献类型:
--
作者:
H. Ben-younes;Rémi Cadène;Nicolas Thome;M. Cord

文献摘要

被引文献

相似文献

多模态表示学习在深度学习社区中越来越受到关注。虽然双线性模型提供了一个有趣的框架来寻找模态的微妙组合,但它们的参数数量随着输入维度的平方增长,使得它们在经典深度学习管道中的实际实现具有挑战性。在本文中,我们介绍了块,一个新的多模态融合的块超对角张量分解的基础上。它利用了块项秩的概念,它概括了张量的秩和模式秩的概念,已经用于多模态融合。它允许定义新的方法来优化融合模型的表达能力和复杂性之间的权衡,并且能够表示模态之间非常精细的交互,同时保持强大的单模态表示。我们通过使用BLOCK来完成两个具有挑战性的任务来展示我们的融合模型的实际意义:视觉问题搜索(VQA)和视觉关系检测(VRD),在这里我们设计了端到端的可学习架构来表示模态之间的相关交互。通过大量的实验,我们表明,块比较有利的VQA和VRD任务的最先进的多模态融合模型。我们的代码可在www.example.com上获得。
Multimodal representation learning is gaining more and more interest within the deep learning community. While bilinear models provide an interesting framework to find subtle combination of modalities, their number of parameters grows quadratically with the input dimensions, making their practical implementation within classical deep learning pipelines challenging. In this paper, we introduce BLOCK, a new multimodal fusion based on the block-superdiagonal tensor decomposition. It leverages the notion of block-term ranks, which generalizes both concepts of rank and mode ranks for tensors, already used for multimodal fusion. It allows to define new ways for optimizing the tradeoff between the expressiveness and complexity of the fusion model, and is able to represent very fine interactions between modalities while maintaining powerful mono-modal representations. We demonstrate the practical interest of our fusion model by using BLOCK for two challenging tasks: Visual Question Answering (VQA) and Visual Relationship Detection (VRD), where we design end-to-end learnable architectures for representing relevant interactions between modalities. Through extensive experiments, we show that BLOCK compares favorably with respect to state-of-the-art multimodal fusion models for both VQA and VRD tasks. Our code is available at https://github.com/Cadene/block.bootstrap.pytorch.