A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others

A Whac-A-Mole Dilemma: Shortcuts Come in Multiples Where Mitigating One Amplifies Others
复制标题

DOI:
10.1109/cvpr52729.2023.01922
复制
发表时间:
2022-12
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Zhiheng Li;I. Evtimov;Albert Gordo;C. Hazirbas;Tal Hassner;Cristian Cantón Ferrer;Chenliang Xu;Mark Ibrahim
Zhiheng Li;I. Evtimov;Albert Gordo;C. Hazirbas;Tal Hassner;Cristian Cantón Ferrer;Chenliang Xu;Mark Ibrahim
中科院分区:
其他
文献类型:
--
作者:
Zhiheng Li;I. Evtimov;Albert Gordo;C. Hazirbas;Tal Hassner;Cristian Cantón Ferrer;Chenliang Xu;Mark Ibrahim

文献摘要

相似文献

机器学习模型已经被发现学习捷径-意想不到的决策规则,无法推广-破坏模型的可靠性。以前的工作解决这个问题的脆弱的假设下,只有一个单一的捷径存在于训练数据。现实世界的图像充斥着从背景到纹理的多种视觉线索。提高视觉系统可靠性的关键是了解现有方法是否可以克服多个捷径或在打地鼠游戏中挣扎,即,减少一条捷径会增加对其他捷径的依赖。为了解决这个缺点,我们提出了两个基准:1)UrbanCars,一个具有精确控制的虚假线索的数据集,以及2)ImageNet-W,一个基于ImageNet的水印评估集,我们发现的一个捷径几乎影响了每一个现代视觉模型。沿着纹理和背景,ImageNet-W允许我们研究自然图像训练中出现的多种快捷方式。我们发现计算机视觉模型,包括大型基础模型--无论训练集、架构和监督如何--在存在多个快捷方式时都会遇到困难。即使是明确设计用于打击捷径的方法,也会陷入打地鼠的困境。为了应对这一挑战,我们提出了最后一层包围,这是一种简单而有效的方法,可以减少多个快捷方式,而不会出现Whac-A-Mole行为。我们的研究结果表明,多捷径缓解是一个被忽视的挑战,对提高视觉系统的可靠性至关重要。数据集和代码已发布:https://github.com/facebookresearch/Whac-A-Mole。
Machine learning models have been found to learn shortcuts—unintended decision rules that are unable to generalize—undermining models' reliability. Previous works address this problem under the tenuous assumption that only a single shortcut exists in the training data. Real-world images are rife with multiple visual cues from background to texture. Key to advancing the reliability of vision systems is understanding whether existing methods can overcome multiple shortcuts or struggle in a Whac-A-Mole game, i.e., where mitigating one shortcut amplifies reliance on others. To address this shortcoming, we propose two benchmarks: 1) UrbanCars, a dataset with precisely controlled spurious cues, and 2) ImageNet-W, an evaluation set based on ImageNet for watermark, a shortcut we discovered affects nearly every modern vision model. Along with texture and background, ImageNet-W allows us to study multiple shortcuts emerging from training on natural images. We find computer vision models, including large foundation models—regardless of training set, architecture, and supervision—struggle when multiple shortcuts are present. Even methods explicitly designed to combat shortcuts struggle in a Whac-A-Mole dilemma. To tackle this challenge, we propose Last Layer Ensemble, a simple-yet-effective method to mitigate multiple shortcuts without Whac-A-Mole behavior. Our results surface multi-shortcut mitigation as an overlooked challenge critical to advancing the reliability of vision systems. The datasets and code are released: https://github.com/facebookresearch/Whac-A-Mole.