Why Fairness Cannot Be Automated: Bridging the Gap Between EU Non-Discrimination Law and AI

Why Fairness Cannot Be Automated: Bridging the Gap Between EU Non-Discrimination Law and AI
复制标题

为什么公平无法自动化:弥合欧盟非歧视法与人工智能之间的差距

DOI:
--
复制
发表时间:
2020
影响因子:
2.9
通讯作者:
Chris Russell
Chris Russell
中科院分区:
法学4区
文献类型:
--
作者:
Sandra Wachter;B. Mittelstadt;Chris Russell

文献摘要

被引文献

相似文献

近年来,出现了大量关于人工智能和机器学习中的偏见、歧视和公平的文献。将这项工作与现有的法律的不歧视框架联系起来,对于创造在不同的法律的制度中实际有用的工具和方法至关重要。虽然从美国法律的角度进行了大量的工作,但相对而言,很少有人绘制欧盟法律的影响和要求。这篇文章解决了算法公平性的法律的、技术的和组织的概念之间的关键差距。通过对欧盟非歧视法以及欧洲法院和各国法院的判例的分析,我们发现欧洲的歧视概念与现有的算法和自动化公平工作之间存在严重的不相容性。无数公平工具包和治理机制中嵌入的公平统计指标与欧洲法院使用的上下文敏感、通常直观且模糊的歧视指标和证据要求之间存在明显差距;我们将这种方法称为“上下文平等”。本文有三个贡献。首先,我们回顾了根据欧盟非歧视法提出索赔的证据要求。由于算法和人类歧视的不同性质,欧盟目前的要求过于上下文,依赖直觉,并对司法解释开放,无法自动化。对提出索赔至关重要的许多概念,如处境不利和弱势群体的构成、所受伤害的严重程度和类型,以及证据的相关性和可接受性要求,都要求司法机构在个案基础上作出规范性或政治性选择。我们表明,在欧洲实现公平或非歧视的自动化可能是不可能的,因为法律在设计上并没有提供一个适合于测试人工智能系统中歧视的静态或同质框架。其次,我们展示了当人工智能而不是人类歧视时,非歧视法提供的法律的保护是如何受到挑战的。人类歧视是由于消极的态度(如陈规定型观念、偏见)和无意的偏见(如组织惯例或内化的陈规定型观念),这可能会向受害者发出歧视已经发生的信号。在算法系统中不存在等效的信令机制和代理。与传统形式的歧视相比,自动化歧视更加抽象和不直观,微妙,无形,难以检测。算法的使用越来越多,破坏了主要依赖直觉的传统法律的补救措施和检测、调查、预防和纠正歧视的程序。一致的评估程序,定义了一个共同的标准,统计证据,以检测和评估初步自动歧视是迫切需要支持法官,监管机构,系统控制器和开发人员,和索赔人。最后,我们研究如何现有的工作在机器学习的公平性排队与程序,根据欧盟非歧视法评估案件。欧洲法院提出了评估表面歧视的“黄金标准”,但尚未转化为自动歧视的标准评估程序。我们提出了“有条件的人口差异”(CDD)作为一个标准的基线统计测量,符合法院的“黄金标准”。为自动化歧视案件建立一套标准的统计证据,有助于确保对涉及人工智能和自动化系统的案件进行一致的评估程序,而不是司法解释。通过这一建议的程序规则性的识别和评估的自动化歧视,我们澄清如何建立考虑到公平的自动化系统尽可能,同时仍然尊重和启用上下文的方法,以司法解释实践根据欧盟非歧视法。
In recent years a substantial literature has emerged concerning bias, discrimination, and fairness in AI and machine learning. Connecting this work to existing legal non-discrimination frameworks is essential to create tools and methods that are practically useful across divergent legal regimes. While much work has been undertaken from an American legal perspective, comparatively little has mapped the effects and requirements of EU law. This Article addresses this critical gap between legal, technical, and organisational notions of algorithmic fairness. Through analysis of EU non-discrimination law and jurisprudence of the European Court of Justice (ECJ) and national courts, we identify a critical incompatibility between European notions of discrimination and existing work on algorithmic and automat-ed fairness. A clear gap exists between statistical measures of fairness as embedded in myriad fairness toolkits and governance mechanisms and the context-sensitive, often intuitive and ambiguous discrimination metrics and evidential requirements used by the ECJ; we refer to this approach as “contextual equality.”This Article makes three contributions. First, we review the evidential requirements to bring a claim under EU non-discrimination law. Due to the disparate nature of algorithmic and human discrimination, the EU’s current requirements are too contextual, reliant on intuition, and open to judicial interpretation to be automated. Many of the concepts fundamental to bringing a claim, such as the composition of the disadvantaged and advantaged group, the severity and type of harm suffered, and requirements for the relevance and admissibility of evidence, require normative or political choices to be made by the judiciary on a case-by-case basis. We show that automating fairness or non-discrimination in Europe may be impossible because the law, by design, does not provide a static or homogenous framework suited to testing for discrimination in AI systems.Second, we show how the legal protection offered by non-discrimination law is challenged when AI, not humans, discriminate. Humans discriminate due to negative attitudes (e.g. stereotypes, prejudice) and unintentional biases (e.g. organisational practices or internalised stereotypes) which can act as a signal to victims that discrimination has occurred. Equivalent signalling mechanisms and agency do not exist in algorithmic systems. Compared to traditional forms of discrimination, automated discrimination is more abstract and unintuitive, subtle, intangible, and difficult to detect. The increasing use of algorithms disrupts traditional legal remedies and procedures for detection, investigation, prevention, and correction of discrimination which have predominantly relied upon intuition. Consistent assessment procedures that define a common standard for statistical evidence to detect and assess prima facie automated discrimination are urgently needed to support judges, regulators, system controllers and developers, and claimants.Finally, we examine how existing work on fairness in machine learning lines up with procedures for assessing cases under EU non-discrimination law. A ‘gold standard’ for assessment of prima facie discrimination has been advanced by the European Court of Justice but not yet translated into standard assessment procedures for automated discrimination. We propose ‘conditional demographic disparity’ (CDD) as a standard baseline statistical measurement that aligns with the Court’s ‘gold standard’. Establishing a standard set of statistical evidence for automated discrimination cases can help ensure consistent procedures for assessment, but not judicial interpretation, of cases involving AI and automated systems. Through this proposal for procedural regularity in the identification and assessment of auto-mated discrimination, we clarify how to build considerations of fairness into automated systems as far as possible while still respecting and enabling the contextual approach to judicial interpretation practiced under EU non-discrimination law.