ProPublica's COMPAS Data Revisited

ProPublica's COMPAS Data Revisited
复制标题

重新审视 ProPublica 的 COMPAS 数据

DOI:
--
复制
发表时间:
2019
期刊:
arXiv.org
影响因子:
--
通讯作者:
M. Barenstein
M. Barenstein
中科院分区:
--
文献类型:
--
作者:
M. Barenstein

文献摘要

被引文献

相似文献

在本文中,我重新审视了ProPublica在2016年收集的COMPAS累犯评分和犯罪历史数据,这些数据在过去三年中引发了“算法公平”或“公平机器学习”这一新兴领域的激烈辩论和研究。ProPublica的COMPAS数据被越来越多的研究用来测试算法公平性的各种定义和方法。本文仔细研究了ProPublica收集的实际数据集。特别是,我检查了被告在COMPAS筛选日期的分布,发现ProPublica在创建其他研究人员最常用的一些关键数据集时犯了一个重要的数据处理错误。具体来说,这些数据集是为了研究最初COMPAS筛查日期后两年内再犯的可能性而建立的。正如我在这篇论文中所展示的,ProPublica在这样的数据集中对累犯实施了两年样本截止规则(而它对非累犯实施了适当的两年样本截止规则)。结果,ProPublica错误地保留了不成比例的累犯。这种数据处理错误导致了有偏见的两年累犯数据集,人为地提高了累犯率。这也影响阳性和阴性预测值。另一方面,这种数据处理错误不会影响ProPublica和其他研究人员强调的一些关键统计指标,例如假阳性和假阴性率,也不会影响整体准确性。
In this paper I re-examine the COMPAS recidivism score and criminal history data collected by ProPublica in 2016, which has fueled intense debate and research in the nascent field of `algorithmic fairness' or `fair machine learning' over the past three years. ProPublica's COMPAS data is used in an ever-increasing number of studies to test various definitions and methodologies of algorithmic fairness. This paper takes a closer look at the actual datasets put together by ProPublica. In particular, I examine the distribution of defendants across COMPAS screening dates and find that ProPublica made an important data processing mistake when it created some of the key datasets most often used by other researchers. Specifically, the datasets built to study the likelihood of recidivism within two years of the original COMPAS screening date. As I show in this paper, ProPublica made a mistake implementing the two-year sample cutoff rule for recidivists in such datasets (whereas it implemented an appropriate two-year sample cutoff rule for non-recidivists). As a result, ProPublica incorrectly kept a disproportionate share of recidivists. This data processing mistake leads to biased two-year recidivism datasets, with artificially high recidivism rates. This also affects the positive and negative predictive values. On the other hand, this data processing mistake does not impact some of the key statistical measures highlighted by ProPublica and other researchers, such as the false positive and false negative rates, nor the overall accuracy.