Fairness without Imputation: A Decision Tree Approach for Fair Prediction with Missing Values

Fairness without Imputation: A Decision Tree Approach for Fair Prediction with Missing Values
复制标题

DOI:
10.1609/aaai.v36i9.21189
复制
发表时间:
2021-09
期刊:
--
影响因子:
--
通讯作者:
Haewon Jeong;Hao Wang;F. Calmon
Haewon Jeong;Hao Wang;F. Calmon
中科院分区:
其他
文献类型:
--
作者:
Haewon Jeong;Hao Wang;F. Calmon

文献摘要

被引文献

相似文献

我们研究了使用缺失值数据训练机器学习模型的公平性问题。尽管文献中有许多公平性干预方法,但大多数都需要完整的训练集作为输入。在实践中,数据可能有缺失值,数据缺失模式可能取决于组属性(例如性别或种族)。简单地将现成的公平学习算法应用于估算数据集可能会导致不公平的模型。在本文中,我们首先从理论上分析了在使用估算数据集进行训练时歧视风险的不同来源。然后,我们提出了一种基于决策树的集成方法,该方法不需要单独的估算和学习过程。相反,我们训练了一棵树,其中包含缺失的属性(MIA),这不需要显式插补,并且我们优化了公平正则化的目标函数。我们证明了我们的方法优于现有的公平干预方法应用于估算数据集,通过对真实世界的数据集的几个实验。
We investigate the fairness concerns of training a machine learning model using data with missing values. Even though there are a number of fairness intervention methods in the literature, most of them require a complete training set as input. In practice, data can have missing values, and data missing patterns can depend on group attributes (e.g. gender or race). Simply applying off-the-shelf fair learning algorithms to an imputed dataset may lead to an unfair model. In this paper, we first theoretically analyze different sources of discrimination risks when training with an imputed dataset. Then, we propose an integrated approach based on decision trees that does not require a separate process of imputation and learning. Instead, we train a tree with missing incorporated as attribute (MIA), which does not require explicit imputation, and we optimize a fairness-regularized objective function. We demonstrate that our approach outperforms existing fairness intervention methods applied to an imputed dataset, through several experiments on real-world datasets.