Quantifying explainable discrimination and removing illegal discrimination in automated decision making

Quantifying explainable discrimination and removing illegal discrimination in automated decision making
复制标题

量化可解释的歧视并消除自动决策中的非法歧视

DOI:
--
复制
发表时间:
2013
影响因子:
2.7
通讯作者:
T. Calders
T. Calders
中科院分区:
计算机科学4区
文献类型:
--
作者:
F. Kamiran;Indrė Žliobaitė;T. Calders

文献摘要

被引文献

相似文献

最近,引入了下面的区分意识分类问题。用于监督学习的历史数据可能包含歧视,例如关于性别的歧视。歧视感知技术所要解决的问题是,在给定敏感属性的情况下,如何在对于给定的敏感属性具有歧视性的历史数据上训练无歧视分类器。处理这一问题的现有技术旨在消除所有歧视,没有考虑到部分歧视可能可以用其他属性来解释。例如,在求职申请中,求职者的教育水平可能是这样一个可以解释的属性。如果数据包含许多受过高等教育的男性候选人,而只有几名受过高等教育的女性,那么男女录取率的差异并不一定反映出性别歧视,因为这可以用不同的教育水平来解释。尽管根据教育水平进行选择会导致更多的男性被录取,但在这种标准上的差异不会被认为是不可取的,也不会被认为是非法的。然而,目前最先进的技术没有考虑到这种性别中立的解释,而且往往反应过度,实际上开始反向歧视,正如我们将在本文中所展示的那样。因此,我们在分类器设计中引入并分析了条件无歧视的精炼概念。我们表明,敏感群体之间的一些决策差异是可以解释的,因此是可以容忍的。因此,我们发展了量化可解释歧视的方法和当一个或多个属性被视为解释性歧视时消除非法歧视的算法技术。对合成和真实世界分类数据集的实验评估表明,新技术在这种新的背景下优于旧技术,因为它们成功地几乎完全消除了不希望看到的歧视,同时保持了可解释的差异不变,允许决策中的差异,只要它们是可解释的。
Recently, the following discrimination-aware classification problem was introduced. Historical data used for supervised learning may contain discrimination, for instance, with respect to gender. The question addressed by discrimination-aware techniques is, given sensitive attribute, how to train discrimination-free classifiers on such historical data that are discriminative, with respect to the given sensitive attribute. Existing techniques that deal with this problem aim at removing all discrimination and do not take into account that part of the discrimination may be explainable by other attributes. For example, in a job application, the education level of a job candidate could be such an explainable attribute. If the data contain many highly educated male candidates and only few highly educated women, a difference in acceptance rates between woman and man does not necessarily reflect gender discrimination, as it could be explained by the different levels of education. Even though selecting on education level would result in more males being accepted, a difference with respect to such a criterion would not be considered to be undesirable, nor illegal. Current state-of-the-art techniques, however, do not take such gender-neutral explanations into account and tend to overreact and actually start reverse discriminating, as we will show in this paper. Therefore, we introduce and analyze the refined notion of conditional non-discrimination in classifier design. We show that some of the differences in decisions across the sensitive groups can be explainable and are hence tolerable. Therefore, we develop methodology for quantifying the explainable discrimination and algorithmic techniques for removing the illegal discrimination when one or more attributes are considered as explanatory. Experimental evaluation on synthetic and real-world classification datasets demonstrates that the new techniques are superior to the old ones in this new context, as they succeed in removing almost exclusively the undesirable discrimination, while leaving the explainable differences unchanged, allowing for differences in decisions as long as they are explainable.