Recovering from Biased Data: Can Fairness Constraints Improve Accuracy?

Recovering from Biased Data: Can Fairness Constraints Improve Accuracy?
复制标题

DOI:
10.4230/lipics.forc.2020.3
复制
发表时间:
2019-12
期刊:
--
影响因子:
--
通讯作者:
Avrim Blum;Kevin Stangl
Avrim Blum;Kevin Stangl
中科院分区:
其他
文献类型:
--
作者:
Avrim Blum;Kevin Stangl

文献摘要

相似文献

文献中已经提出了多个公平性约束,其动机是对机器学习分类器如何不公平地对待人口统计学群体的一系列担忧。在这项工作中,我们考虑了不同的动机;从有偏见的训练数据中学习。我们分析了训练数据可能存在偏见的几种方式,包括对弱势群体成员进行更具噪音或负面偏见的标记过程,或者弱势群体中积极或消极示例的流行率降低,或者两者兼而有之。给定这种有偏差的训练数据,经验风险最小化(ERM)可能会产生一个分类器,不仅是有偏差的,而且在真实数据分布上具有次优的准确性。我们研究的能力,公平约束的ERM纠正这个问题。特别是,我们发现平等机会公平约束(Hardt,Price和Srebro 2016)与ERM相结合将可证明在一系列偏差模型下恢复贝叶斯最优分类器。我们还考虑了其他恢复方法,包括重新加权训练数据,均衡赔率和人口统计学平价。这些理论结果提供了额外的动机,考虑公平的干预措施,即使演员主要关心的准确性。
Multiple fairness constraints have been proposed in the literature, motivated by a range of concerns about how demographic groups might be treated unfairly by machine learning classifiers. In this work we consider a different motivation; learning from biased training data. We posit several ways in which training data may be biased, including having a more noisy or negatively biased labeling process on members of a disadvantaged group, or a decreased prevalence of positive or negative examples from the disadvantaged group, or both. Given such biased training data, Empirical Risk Minimization (ERM) may produce a classifier that not only is biased but also has suboptimal accuracy on the true data distribution. We examine the ability of fairness-constrained ERM to correct this problem. In particular, we find that the Equal Opportunity fairness constraint (Hardt, Price, and Srebro 2016) combined with ERM will provably recover the Bayes Optimal Classifier under a range of bias models. We also consider other recovery methods including reweighting the training data, Equalized Odds, and Demographic Parity. These theoretical results provide additional motivation for considering fairness interventions even if an actor cares primarily about accuracy.