Domain-robust VQA with diverse datasets and methods but no target labels

Domain-robust VQA with diverse datasets and methods but no target labels
复制标题

DOI:
10.1109/cvpr46437.2021.00697
复制
发表时间:
2021-03
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Mingda Zhang;Tristan D. Maidment;Ahmad Diab;Adriana Kovashka;R. Hwa
Mingda Zhang;Tristan D. Maidment;Ahmad Diab;Adriana Kovashka;R. Hwa
中科院分区:
其他
文献类型:
--
作者:
Mingda Zhang;Tristan D. Maidment;Ahmad Diab;Adriana Kovashka;R. Hwa

文献摘要

被引文献

相似文献

计算机视觉方法过度适应数据集细节的观察激发了各种尝试,使对象识别模型对领域变化具有鲁棒性。然而,关于领域鲁棒性视觉问答方法的类似工作非常有限。由于额外的复杂性,VQA 的域适应与对象识别的适应不同:VQA 模型处理多模态输入,方法包含具有不同模块的多个步骤,导致复杂的优化,并且不同数据集中的答案空间差异很大。为了应对这些挑战,我们首先在视觉和文本空间中量化流行的 VQA 数据集之间的域转移。为了理清由不同模态产生的数据集之间的变化,我们还分别在图像和问题域中构建合成变化。其次,我们测试了不同系列的 VQA 方法(经典的双流、变压器和神经符号方法)对这些转变的稳健性。第三,我们测试现有域适应方法的适用性,并设计一种新的方法来弥合 VQA 域差距,并根据特定的 VQA 模型进行调整。为了模拟现实世界泛化的设置,我们专注于无监督域适应和开放式分类任务制定。
The observation that computer vision methods overfit to dataset specifics has inspired diverse attempts to make object recognition models robust to domain shifts. However, similar work on domain-robust visual question answering methods is very limited. Domain adaptation for VQA differs from adaptation for object recognition due to additional complexity: VQA models handle multimodal inputs, methods contain multiple steps with diverse modules resulting in complex optimization, and answer spaces in different datasets are vastly different. To tackle these challenges, we first quantify domain shifts between popular VQA datasets, in both visual and textual space. To disentangle shifts between datasets arising from different modalities, we also construct synthetic shifts in the image and question domains separately. Second, we test the robustness of different families of VQA methods (classic two-stream, transformer, and neuro-symbolic methods) to these shifts. Third, we test the applicability of existing domain adaptation methods and devise a new one to bridge VQA domain gaps, adjusted to specific VQA models. To emulate the setting of real-world generalization, we focus on unsupervised domain adaptation and the open-ended classification task formulation.