Using sensitive personal data may be necessary for avoiding discrimination in data-driven decision models

Using sensitive personal data may be necessary for avoiding discrimination in data-driven decision models
复制标题

为了避免数据驱动的决策模型中的歧视,可能需要使用敏感的个人数据

DOI:
--
复制
发表时间:
2016
影响因子:
4.1
通讯作者:
B. Custers
B. Custers
中科院分区:
计算机科学2区
文献类型:
--
作者:
Indrė Žliobaitė;B. Custers

文献摘要

参考文献

被引文献

相似文献

越来越多的日常生活决策是使用算法做出的。我们所说的算法是指使用数据挖掘从历史数据中捕获的预测模型(决策规则)。这样的模特通常决定我们支付的价格,选择我们在网上看到的广告和新闻,匹配工作描述和候选人简历,决定谁获得贷款,谁接受额外的机场安检,或者谁获得假释。然而,越来越多的证据表明,算法的决策可能会歧视人,即使计算过程是公平和善意的。这是由于学习数据有偏差或不具代表性,再加上无意的建模过程造成的。从监管角度来看,这个问题有两种趋势:(1)确保数据驱动的决策不具有歧视性;(2)将私人数据的总体收集和存储限制在必要的最低限度。本文表明,从计算的角度来看,这两个目标是矛盾的。我们用标准回归模型从经验和理论上证明,为了确保决策模型是非歧视性的,例如关于种族的决策模型,需要在模型建立过程中使用敏感的种族信息。当然,在模型准备好之后,种族不应该被要求作为决策的输入变量。从监管角度来看,这有一个重要的含义:为了保证算法的公平性,收集敏感的个人数据是必要的,立法需要找到合理的方法,允许在建模过程中使用这些数据。
Increasing numbers of decisions about everyday life are made using algorithms. By algorithms we mean predictive models (decision rules) captured from historical data using data mining. Such models often decide prices we pay, select ads we see and news we read online, match job descriptions and candidate CVs, decide who gets a loan, who goes through an extra airport security check, or who gets released on parole. Yet growing evidence suggests that decision making by algorithms may discriminate people, even if the computing process is fair and well-intentioned. This happens due to biased or non-representative learning data in combination with inadvertent modeling procedures. From the regulatory perspective there are two tendencies in relation to this issue: (1) to ensure that data-driven decision making is not discriminatory, and (2) to restrict overall collecting and storing of private data to a necessary minimum. This paper shows that from the computing perspective these two goals are contradictory. We demonstrate empirically and theoretically with standard regression models that in order to make sure that decision models are non-discriminatory, for instance, with respect to race, the sensitive racial information needs to be used in the model building process. Of course, after the model is ready, race should not be required as an input variable for decision making. From the regulatory perspective this has an important implication: collecting sensitive personal data is necessary in order to guarantee fairness of algorithms, and law making needs to find sensible ways to allow using such data in the modeling process.
自动化作为反偏见干预的悖论
DOI: --
发表时间: 2020
期刊: Cardozo law review
影响因子: --
作者:
Ajunwa, Ifeoma
通讯作者: Ajunwa, Ifeoma