On the Direction of Discrimination: An Information-Theoretic Analysis of Disparate Impact in Machine Learning

On the Direction of Discrimination: An Information-Theoretic Analysis of Disparate Impact in Machine Learning
复制标题

关于歧视的方向:机器学习中不同影响的信息论分析

DOI:
--
复制
发表时间:
2018
期刊:
International Symposium on Information Theory
影响因子:
--
通讯作者:
F. Calmon
F. Calmon
中科院分区:
--
文献类型:
--
作者:
Hao Wang;Berk Ustun;F. Calmon

文献摘要

被引文献

相似文献

在机器学习的背景下,不同的影响是指一种系统歧视的形式,模型的输出分布取决于敏感属性(例如种族或性别)的值。在本文中,我们提出了一个信息论框架来分析二元分类模型的不同影响。我们将该模型视为固定渠道,并将不同的影响量化为两组输出分布的差异。我们的目标是找到一个校正函数,可以扰乱每个组的输入分布以对齐其输出分布。我们提出了一个优化问题,可以解决该问题以获得校正函数,该函数将使输出分布在统计上无法区分。我们推导出封闭式表达式来有效计算校正函数,并展示了我们的框架在基于 ProPublica COMPAS 数据集的累犯预测问题上的优势。
In the context of machine learning, disparate impact refers to a form of systematic discrimination whereby the output distribution of a model depends on the value of a sensitive attribute (e.g., race or gender). In this paper, we propose an information-theoretic framework to analyze the disparate impact of a binary classification model. We view the model as a fixed channel, and quantify disparate impact as the divergence in output distributions over two groups. Our aim is to find a correction function that can perturb the input distributions of each group to align their output distributions. We present an optimization problem that can be solved to obtain a correction function that will make the output distributions statistically indistinguishable. We derive closed-form expressions to efficiently compute the correction function, and demonstrate the benefits of our framework on a recidivism prediction problem based on the ProPublica COMPAS dataset.