SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question Answering

SwapMix: Diagnosing and Regularizing the Over-Reliance on Visual Context in Visual Question Answering
复制标题

DOI:
10.1109/cvpr52688.2022.00502
复制
发表时间:
2022-04
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Vipul Gupta;Zhuowan Li;Adam Kortylewski;Chenyu Zhang-;Yingwei Li;A. Yuille
Vipul Gupta;Zhuowan Li;Adam Kortylewski;Chenyu Zhang-;Yingwei Li;A. Yuille
中科院分区:
其他
文献类型:
--
作者:
Vipul Gupta;Zhuowan Li;Adam Kortylewski;Chenyu Zhang-;Yingwei Li;A. Yuille

文献摘要

被引文献

相似文献

虽然视觉问答 (VQA) 发展迅速,但之前的工作引起了人们对当前 VQA 模型稳健性的担忧。在这项工作中,我们从一个新颖的角度研究 VQA 模型的稳健性:视觉上下文。我们认为模型过度依赖视觉上下文,即图像中不相关的对象来进行预测。为了诊断模型对视觉上下文的依赖并测量其鲁棒性,我们提出了一种简单而有效的扰动技术,SwapMix。 SwapMix 通过将不相关上下文对象的特征与数据集中其他对象的特征交换来扰乱视觉上下文。使用 SwapMix,我们能够更改代表性 VQA 模型超过 45% 的问题的答案。此外,我们用完美的视觉训练模型,发现上下文的过度依赖很大程度上取决于视觉表示的质量。除了诊断之外,SwapMix 还可以在训练期间用作数据增强策略,以规范上下文过度依赖。通过交换上下文对象特征,可以有效抑制模型对上下文的依赖。使用 SwapMix 研究了两个代表性的 VQA 模型:共同注意力模型 MCAN 和大规模预训练模型 LXMERT。我们在流行的 GQA 数据集上进行的实验表明了 SwapMix 在诊断模型稳健性和规范对视觉上下文的过度依赖方面的有效性。我们方法的代码可在 https://github.com/vipulgupta1011/swapmix 获取
While Visual Question Answering (VQA) has progressed rapidly, previous works raise concerns about robustness of current VQA models. In this work, we study the robustness of VQA models from a novel perspective: visual context. We suggest that the models over-rely on the visual context, i.e., irrelevant objects in the image, to make predictions. To diagnose the models' reliance on visual context and measure their robustness, we propose a simple yet effective perturbation technique, SwapMix. SwapMix perturbs the visual context by swapping features of irrelevant context objects with features from other objects in the dataset. Using SwapMix we are able to change answers to more than 45% of the questions for a representative VQA model. Additionally, we train the models with perfect sight and find that the context over-reliance highly depends on the quality of visual representations. In addition to diagnosing, SwapMix can also be applied as a data augmentation strategy during training in order to regularize the context over-reliance. By swapping the context object features, the model reliance on context can be suppressed effectively. Two representative VQA models are studied using SwapMix: a co-attention model MCAN and a large-scale pretrained model LXMERT. Our experiments on the popular GQA dataset show the effectiveness of SwapMix for both diagnosing model robustness, and regularizing the over-reliance on visual context. The code for our method is available at https://github.com/vipulgupta1011/swapmix