Understanding Instance-Level Impact of Fairness Constraints

Understanding Instance-Level Impact of Fairness Constraints
复制标题

DOI:
10.48550/arxiv.2206.15437
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Jialu Wang;X. Wang;Yang Liu
Jialu Wang;X. Wang;Yang Liu
中科院分区:
其他
文献类型:
--
作者:
Jialu Wang;X. Wang;Yang Liu

文献摘要

相似文献

在文献中已经提出了各种公平性约束,以减轻组级统计偏差。它们的影响主要是针对与种族或性别等一系列敏感属性相对应的不同人口群体进行评估的。尽管如此,社区尚未观察到对在实例级别施加公平约束的效果如何进行足够的探索。基于影响函数的概念,一个衡量训练样本对目标模型及其预测性能的影响的指标,这项工作研究了在施加公平性约束时训练样本的影响。我们发现,在一定的假设下,公平性约束的影响函数可以被分解为训练样本的核化组合。提出的公平性影响函数的一个有希望的应用是通过对它们的影响分数进行排名来识别可能引起模型歧视的可疑训练示例。我们通过大量的实验证明,在重要数据示例的子集上进行训练会导致较低的公平性违规,同时会牺牲准确性。
A variety of fairness constraints have been proposed in the literature to mitigate group-level statistical bias. Their impacts have been largely evaluated for different groups of populations corresponding to a set of sensitive attributes, such as race or gender. Nonetheless, the community has not observed sufficient explorations for how imposing fairness constraints fare at an instance level. Building on the concept of influence function, a measure that characterizes the impact of a training example on the target model and its predictive performance, this work studies the influence of training examples when fairness constraints are imposed. We find out that under certain assumptions, the influence function with respect to fairness constraints can be decomposed into a kernelized combination of training examples. One promising application of the proposed fairness influence function is to identify suspicious training examples that may cause model discrimination by ranking their influence scores. We demonstrate with extensive experiments that training on a subset of weighty data examples leads to lower fairness violations with a trade-off of accuracy.