Fair Regression under Sample Selection Bias

Fair Regression under Sample Selection Bias
复制标题

样本选择偏差下的公平回归

DOI:
10.1109/bigdata55660.2022.10021107
复制
发表时间:
2022
期刊:
2022 IEEE International Conference on Big Data (Big Data
影响因子:
--
通讯作者:
Tong, Hanghang
Tong, Hanghang
中科院分区:
--
文献类型:
--
作者:
Du, Wei;Wu, Xintao;Tong, Hanghang

文献摘要

参考文献

被引文献

相似文献

近年来,公平回归的研究主要集中在发展新的公平概念和近似方法,因为目标变量甚至敏感属性在回归设置中是连续的。然而,所有以前的公平回归研究都假设训练数据和测试数据来自相同的分布。在真实的世界中,由于训练数据和测试数据之间的样本选择偏差,这一假设经常被违反。在本文中,我们开发了一个框架下的样本选择偏差的公平回归时,从训练数据的一组样本的因变量值丢失作为另一个隐藏的过程的结果。我们的框架采用了经典的Heckman模型的偏差校正和拉格朗日对偶,以实现公平的回归基于各种公平的概念。Heckman模型描述了样本选择过程,并使用称为逆米尔斯比(IMR)的衍生变量来校正样本选择偏倚。我们使用公平不等式和等式约束来描述各种公平概念,并应用拉格朗日对偶理论将原问题转化为对偶凸优化问题。对于两个流行的公平性概念,平均差和均方误差差,我们推导出明确的公式,无需迭代优化,和皮尔逊相关,我们推导出它的条件,实现强对偶。我们在三个真实世界的数据集上进行了实验,实验结果证明了该方法在效用和公平性指标方面的有效性。
Recent research on fair regression focused on developing new fairness notions and approximation methods as target variables and even the sensitive attribute are continuous in the regression setting. However, all previous fair regression research assumed the training data and testing data are drawn from the same distributions. This assumption is often violated in real world due to the sample selection bias between the training and testing data. In this paper, we develop a framework for fair regression under sample selection bias when dependent variable values of a set of samples from the training data are missing as a result of another hidden process. Our framework adopts the classic Heckman model for bias correction and the Lagrange duality to achieve fairness in regression based on a variety of fairness notions. Heckman model describes the sample selection process and uses a derived variable called the Inverse Mills Ratio (IMR) to correct sample selection bias. We use fairness inequality and equality constraints to describe a variety of fairness notions and apply the Lagrange duality theory to transform the primal problem into the dual convex optimization. For the two popular fairness notions, mean difference and mean squared error difference, we derive explicit formulas without iterative optimization, and for Pearson correlation, we derive its conditions of achieving strong duality. We conduct experiments on three real-world datasets and the experimental results demonstrate the approach’s effectiveness in terms of both utility and fairness metrics.
具有公平性约束的回归的非凸优化
DOI: --
发表时间: 2018
期刊: The 35th International Conference on Machine Learning (ICML2018)
影响因子: --
作者:
Junpei Komiyama;Akiko Takeda;Junya Honda;Hajime Shimao
通讯作者: Hajime Shimao
排名和回归的成对公平性
DOI: 10.1609/aaai.v34i04.5970
发表时间: 2019
期刊: ArXiv
影响因子: --
作者:
H. Narasimhan;Andrew Cotter;Maya R. Gupta;S. Wang
通讯作者: S. Wang
通过插件估计器进行公平回归并通过统计保证进行重新校准
DOI: --
发表时间: 2020
期刊: Neural Information Processing Systems
影响因子: --
作者:
Evgenii Chzhen;Christophe Denis;Mohamed Hebiri;L. Oneto;M. Pontil
通讯作者: M. Pontil
DOI: 10.1609/aaai.v34i04.6002
发表时间: 2019-03
期刊: --
影响因子: --
作者:
Ashkan Rezaei;Rizal Fathony;Omid Memarrast;Brian D. Ziebart
通讯作者: Ashkan Rezaei;Rizal Fathony;Omid Memarrast;Brian D. Ziebart
公平机器学习中来自偏见数据的残余不公平
DOI: --
发表时间: 2018
期刊: Proceedings of the 35th International Conference on Machine Learning
影响因子: --
作者:
Kallus, Nathan;Zhou, Angela
通讯作者: Zhou, Angela