Improving Visual Grounding by Encouraging Consistent Gradient-Based Explanations

Improving Visual Grounding by Encouraging Consistent Gradient-Based Explanations
复制标题

DOI:
10.1109/cvpr52729.2023.01837
复制
发表时间:
2022-06
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Ziyan Yang;Kushal Kafle;Franck Dernoncourt;Vicente Ord'onez Rom'an
Ziyan Yang;Kushal Kafle;Franck Dernoncourt;Vicente Ord'onez Rom'an
中科院分区:
其他
文献类型:
--
作者:
Ziyan Yang;Kushal Kafle;Franck Dernoncourt;Vicente Ord'onez Rom'an

文献摘要

相似文献

我们提出了一种基于边际的损失来调整联合视觉语言模型,以便它们基于梯度的解释与人类为相对较小的基础数据集提供的区域级注释一致。我们将这一目标称为注意掩码一致性(AMC),并证明了它比以往依赖视觉语言模型对对象检测器的输出进行评分的方法产生更好的视觉基础效果。特别是,在标准视觉语言建模目标的基础上与AMC一起训练的模型在Flickr30k视觉接地基准中获得了86.49%的最新准确率,与在相同监督水平下训练的最佳先前模型相比,绝对提高了5.38%。我们的方法在已建立的指代表达理解基准上也表现得非常好,在RefCOCO+的简单测试中获得了80.34%的准确率,在困难的分裂中获得了64.55%的准确率。AMC是有效的,易于实现,并且是通用的,因为它可以被任何视觉语言模型采用,并且可以使用任何类型的区域注释。
We propose a margin-based loss for tuning joint vision-language models so that their gradient-based explanations are consistent with region-level annotations provided by humans for relatively smaller grounding datasets. We refer to this objective as Attention Mask Consistency (AMC) and demonstrate that it produces superior visual grounding results than previous methods that rely on using vision-language models to score the outputs of object detectors. Particularly, a model trained with AMC on top of standard vision-language modeling objectives obtains a state-of-the-art accuracy of 86.49% in the Flickr30k visual grounding benchmark, an absolute improvement of 5.38% when compared to the best previous model trained under the same level of supervision. Our approach also performs exceedingly well on established benchmarks for referring expression comprehension where it obtains 80.34% accuracy in the easy test of RefCOCO+, and 64.55% in the difficult split. AMC is effective, easy to implement, and is general as it can be adopted by any vision-language model, and can use any type of region annotations.