Bayes-Optimal Fair Classification with Linear Disparity Constraints via Pre-, In-, and Post-processing

Bayes-Optimal Fair Classification with Linear Disparity Constraints via Pre-, In-, and Post-processing
复制标题

通过预处理、中处理和后处理实现具有线性视差约束的贝叶斯最优公平分类

DOI:
--
复制
发表时间:
2024
期刊:
arXiv.org
影响因子:
--
通讯作者:
Edgar Dobriban
Edgar Dobriban
中科院分区:
--
文献类型:
--
作者:
Xianli Zeng;Guang Cheng;Edgar Dobriban

文献摘要

被引文献

相似文献

机器学习算法可能对受保护群体产生不同的影响。为了解决这个问题,我们开发了贝叶斯最优公平分类方法,旨在最大限度地减少给定群体公平约束下的分类错误。我们引入 emph{线性差异度量} 的概念,它是概率分类器的线性函数;和 emph{双线性差异度量},它们在分组回归函数中也是线性的。我们证明了几种流行的不平等衡量标准——人口平等、机会平等和预测平等的偏差——是双线性的。通过揭示与内曼-皮尔逊引理的联系,我们在单个线性差异度量下找到了贝叶斯最优公平分类器的形式。对于双线性差异度量,贝叶斯最优公平分类器成为分组阈值规则。我们的方法还可以处理多个公平性约束(例如均等赔率),以及在预测阶段无法使用受保护属性时的常见场景。利用我们的理论结果,我们设计了在双线性视差约束下学习公平贝叶斯最优分类器的方法。我们的方法涵盖了三种流行的公平感知分类方法,即通过预处理(公平上采样和下采样)、处理中(公平成本敏感分类)和后处理(公平插件规则)。我们的方法直接控制差异,同时实现近乎最优的公平性与准确性权衡。我们凭经验表明,我们的方法与现有算法相比具有优势。
Machine learning algorithms may have disparate impacts on protected groups. To address this, we develop methods for Bayes-optimal fair classification, aiming to minimize classification error subject to given group fairness constraints. We introduce the notion of emph{linear disparity measures}, which are linear functions of a probabilistic classifier; and emph{bilinear disparity measures}, which are also linear in the group-wise regression functions. We show that several popular disparity measures -- the deviations from demographic parity, equality of opportunity, and predictive equality -- are bilinear. We find the form of Bayes-optimal fair classifiers under a single linear disparity measure, by uncovering a connection with the Neyman-Pearson lemma. For bilinear disparity measures, Bayes-optimal fair classifiers become group-wise thresholding rules. Our approach can also handle multiple fairness constraints (such as equalized odds), and the common scenario when the protected attribute cannot be used at the prediction phase. Leveraging our theoretical results, we design methods that learn fair Bayes-optimal classifiers under bilinear disparity constraints. Our methods cover three popular approaches to fairness-aware classification, via pre-processing (Fair Up- and Down-Sampling), in-processing (Fair Cost-Sensitive Classification) and post-processing (a Fair Plug-In Rule). Our methods control disparity directly while achieving near-optimal fairness-accuracy tradeoffs. We show empirically that our methods compare favorably to existing algorithms.