Bias Mimicking: A Simple Sampling Approach for Bias Mitigation

Bias Mimicking: A Simple Sampling Approach for Bias Mitigation
复制标题

DOI:
10.1109/cvpr52729.2023.01945
复制
发表时间:
2022-09
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Maan Qraitem;Kate Saenko;Bryan A. Plummer
Maan Qraitem;Kate Saenko;Bryan A. Plummer
中科院分区:
其他
文献类型:
--
作者:
Maan Qraitem;Kate Saenko;Bryan A. Plummer

文献摘要

相似文献

先前的研究表明,视觉识别数据集经常在类别标签$Y$(例如程序员)内低估偏见组$B$(例如女性)。这种数据集偏差会导致模型学习到类标签和偏差组(如年龄、性别或种族)之间的虚假相关性。解决此问题的大多数最新方法都需要对体系结构进行重大更改,或者需要更多超参数调优的附加损失函数。另外,来自类不平衡文献的数据采样基线(例如,Undersampling, Upweighting)通常可以在一行代码中实现,并且通常没有超参数,提供了更便宜和更有效的解决方案。然而,这些方法都有明显的缺点。例如,欠采样会使每个历元的输入分布下降很大一部分,而过采样会重复采样,导致过拟合。为了解决这些缺点,我们引入了一种新的类条件采样方法:Bias miming。该方法基于这样的观察:如果一个类$c$偏差分布,即$P_{D}(B\vert Y=c)$在每个$c^{\prime}\neq c$上被模拟,那么$Y$和$B$在统计上是独立的。利用这一概念,BM通过一种新颖的训练过程,确保模型暴露于每个历元的整个分布,而不重复样本。因此,在四个基准中,Bias mimics将代表性不足的群体的抽样方法的准确性提高了3%,同时保持并有时提高了非抽样方法的性能。代码:https://github.com/mqraitem/Bias-Mimicking
Prior work has shown that Visual Recognition datasets frequently underrepresent bias groups $B$ (e.g. Female) within class labels $Y$ (e.g. Programmers). This dataset bias can lead to models that learn spurious correlations between class labels and bias groups such as age, gender, or race. Most recent methods that address this problem require significant architectural changes or additional loss functions requiring more hyper-parameter tuning. Alternatively, data sampling baselines from the class imbalance literature (e.g. Undersampling, Upweighting), which can often be implemented in a single line of code and often have no hyperparameters, offer a cheaper and more efficient solution. However, these methods suffer from significant shortcomings. For example, Undersampling drops a significant part of the input distribution per epoch while Oversampling repeats samples, causing overfitting. To address these shortcomings, we introduce a new class-conditioned sampling method: Bias Mimicking. The method is based on the observation that if a class $c$ bias distribution, i.e. $P_{D}(B\vert Y=c)$ is mimicked across every $c^{\prime}\neq c$, then $Y$ and $B$ are statistically independent. Using this notion, BM, through a novel training procedure, ensures that the model is exposed to the entire distribution per epoch without repeating samples. Consequently, Bias Mimicking improves underrepresented groups' accuracy of sampling methods by 3% over four benchmarks while maintaining and sometimes improving performance over nonsampling methods. Code: https://github.com/mqraitem/Bias-Mimicking