Data mining for discrimination discovery

Data mining for discrimination discovery
复制标题

用于发现歧视的数据挖掘

DOI:
--
复制
发表时间:
2010
期刊:
TKDD
影响因子:
--
通讯作者:
F. Turini
F. Turini
中科院分区:
--
文献类型:
--
作者:
S. Ruggieri;D. Pedreschi;F. Turini

文献摘要

被引文献

相似文献

在民权法中,歧视是指基于某一类别或少数群体的成员身份而不考虑个人价值而对人的不公平或不平等待遇。经济学和人文科学的研究人员对信贷、抵押贷款、保险、劳动力市场和教育方面的歧视进行了调查。随着自动决策支持系统的出现,如信用评分系统,数据收集的简便性给数据分析人员在打击歧视方面带来了几个挑战。在本文中,我们介绍了通过数据挖掘在由人工或自动系统获取的历史决策记录的数据集中发现歧视的问题。我们通过建模受法律保护的群体和上下文来形式化直接和间接歧视发现的过程,其中歧视发生在基于分类规则的语法中。基本上,从数据集中提取的分类规则允许揭示非法歧视的背景,其中受法律保护的群体的负担程度是通过扩展分类规则的提升措施来正式确定的。在直接歧视中,可以直接挖掘提取的规则来搜索歧视上下文。在间接歧视方面,挖掘过程需要一些背景知识作为进一步的投入,例如,人口普查数据,与提取的规则相结合,可能会揭示歧视性决定的背景。将提取的分类规则与背景知识相结合所采用的策略称为推理模型。在本文中,我们提出了两种推理模型,并提供了实现它们的自动过程。我们的结果在德国信贷数据集和PKDD发现挑战1999年财务数据集上提供了经验评估。
In the context of civil rights law, discrimination refers to unfair or unequal treatment of people based on membership to a category or a minority, without regard to individual merit. Discrimination in credit, mortgage, insurance, labor market, and education has been investigated by researchers in economics and human sciences. With the advent of automatic decision support systems, such as credit scoring systems, the ease of data collection opens several challenges to data analysts for the fight against discrimination. In this article, we introduce the problem of discovering discrimination through data mining in a dataset of historical decision records, taken by humans or by automatic systems. We formalize the processes of direct and indirect discrimination discovery by modelling protected-by-law groups and contexts where discrimination occurs in a classification rule based syntax. Basically, classification rules extracted from the dataset allow for unveiling contexts of unlawful discrimination, where the degree of burden over protected-by-law groups is formalized by an extension of the lift measure of a classification rule. In direct discrimination, the extracted rules can be directly mined in search of discriminatory contexts. In indirect discrimination, the mining process needs some background knowledge as a further input, for example, census data, that combined with the extracted rules might allow for unveiling contexts of discriminatory decisions. A strategy adopted for combining extracted classification rules with background knowledge is called an inference model. In this article, we propose two inference models and provide automatic procedures for their implementation. An empirical assessment of our results is provided on the German credit dataset and on the PKDD Discovery Challenge 1999 financial dataset.