Big Data's Disparate Impact

Big Data's Disparate Impact
复制标题

DOI:
10.15779/z38bg31
复制
发表时间:
2016-06-01
影响因子:
2.4
通讯作者:
Selbst, Andrew D.
Selbst, Andrew D.
中科院分区:
法学2区
文献类型:
--
作者:
Barocas, Solon;Selbst, Andrew D.

文献摘要

被引文献

相似文献

数据挖掘等算法技术的支持者认为,这些技术可以消除决策过程中的人类偏见。但算法的好坏取决于它所处理的数据。数据往往是不完美的,这使得这些算法继承了先前决策者的偏见。在其他情况下,数据可能只是反映了整个社会中普遍存在的偏见。在另一些情况下,数据挖掘可以发现令人惊讶的有用规律,这些规律实际上只是预先存在的排斥性和不平等模式。对数据挖掘的盲目依赖可能会使历史上处于不利地位的弱势群体无法充分参与社会。更糟糕的是,由于由此产生的歧视几乎总是算法使用过程中无意中出现的属性,而不是程序员有意识的选择,因此很难确定问题的根源,也很难向法院解释。本文通过美国反歧视法的视角来考察这些问题,特别是通过第七章禁止就业歧视。在缺乏明显的歧视意图的情况下,对数据挖掘的受害者来说,最好的理论希望似乎在于差异影响理论。然而,判例法和平等就业机会委员会(Equal Employment Opportunity Commission)的统一指南(Uniform Guidelines)认为,当一项实践的结果可以预测未来的就业结果时,它就可以被证明是一种商业必要性,而数据挖掘就是专门为发现这种统计相关性而设计的。除非有一种合理可行的方法来证明这些发现是虚假的,否则第七章似乎会为其使用提供支持,即使它发现的相关性往往反映了偏见的历史模式,他人对受保护群体成员的歧视,或者基础数据中的缺陷。解决这种非故意歧视的根源并弥补相应的法律缺陷在技术上、法律上和政治上都是困难的。计算所能完成的工作有许多实际的限制。例如,当由于被挖掘的数据本身是过去故意歧视的结果而产生歧视时,通常没有明显的方法来调整历史数据以消除这种污染。在数据挖掘完成后改变数据挖掘结果的纠正措施将涉足法律和政治上有争议的领域。这些改革面临的挑战凸显了反歧视法背后的两大理论之间的紧张关系:反分类和反从属。要找到解决大数据差异性影响的办法,需要的不仅仅是尽最大努力消除偏见;这需要对“歧视”和“公平”的含义进行全面的重新审视。
Advocates of algorithmic techniques like data mining argue that these techniques eliminate human biases from the decision-making process. But an algorithm is only as good as the data it works with. Data is frequently imperfect in ways that allow these algorithms to inherit the prejudices of prior decision makers. In other cases, data may simply reflect the widespread biases that persist in society at large. In still others, data mining can discover surprisingly useful regularities that are really just preexisting patterns of exclusion and inequality. Unthinking reliance on data mining can deny historically disadvantaged and vulnerable groups full participation in society. Worse still, because the resulting discrimination is almost always an unintentional emergent property of the algorithm's use rather than a conscious choice by its programmers, it can be unusually hard to identify the source of the problem or to explain it to a court.This Essay examines these concerns through the lens of American antidiscrimination law-more particularly, through Title VII's prohibition of discrimination in employment. In the absence of a demonstrable intent to discriminate, the best doctrinal hope for data mining's victims would seem to lie in disparate impact doctrine. Case law and the Equal Employment Opportunity Commission's Uniform Guidelines, though, hold that a practice can be justified as a business necessity when its outcomes are predictive of future employment outcomes, and data mining is specifically designed to find such statistical correlations. Unless there is a reasonably practical way to demonstrate that these discoveries are spurious, Title VII would appear to bless its use, even though the correlations it discovers will often reflect historic patterns of prejudice, others' discrimination against members of protected groups, or flaws in the underlying data.Addressing the sources of this unintentional discrimination and remedying the corresponding deficiencies in the law will be difficult technically, difficult legally, and difficult politically. There are a number of practical limits to what can be accomplished computationally. For example, when discrimination occurs because the data being mined is itself a result of past intentional discrimination, there is frequently no obvious method to adjust historical data to rid it of this taint. Corrective measures that alter the results of the data mining after it is complete would tread on legally and politically disputed terrain. These challenges for reform throw into stark relief the tension between the two major theories underlying antidiscrimination law: anticlassification and antisubordination. Finding a solution to big data's disparate impact will require more than best efforts to stamp out prejudice and bias; it will require a wholesale reexamination of the meanings of "discrimination" and "fairness."