Discrimination-aware data mining

Discrimination-aware data mining
复制标题

DOI:
10.1145/1401890.1401959
复制
发表时间:
2008-08
期刊:
ArXiv
影响因子:
--
通讯作者:
D. Pedreschi;S. Ruggieri;F. Turini
D. Pedreschi;S. Ruggieri;F. Turini
中科院分区:
其他
文献类型:
--
作者:
D. Pedreschi;S. Ruggieri;F. Turini

文献摘要

被引文献

相似文献

在民权法的背景下,歧视是指基于成员资格或少数派的成员资格对人的不公平或不平等的待遇,而无需考虑个人优点。通过数据挖掘技术从数据库中提取的规则(例如分类或关联规则)在用于福利或信用批准等决策任务时,可以在上述意义上具有歧视性。在本文中,引入和研究了歧视性分类规则的概念。提供非歧视的保证被证明是一项非琐碎的任务。一种天真的方法,例如夺走所有歧视性属性,在其他背景知识可用时被证明是不够的。我们的方法可以通过背景知识与明显的歧视性规则与歧视性规则和歧视性规则相关的正式结果进行精确表述。还提供了对德国信用数据集结果的经验评估。
In the context of civil rights law, discrimination refers to unfair or unequal treatment of people based on membership to a category or a minority, without regard to individual merit. Rules extracted from databases by data mining techniques, such as classification or association rules, when used for decision tasks such as benefit or credit approval, can be discriminatory in the above sense. In this paper, the notion of discriminatory classification rules is introduced and studied. Providing a guarantee of non-discrimination is shown to be a non trivial task. A naive approach, like taking away all discriminatory attributes, is shown to be not enough when other background knowledge is available. Our approach leads to a precise formulation of the redlining problem along with a formal result relating discriminatory rules with apparently safe ones by means of background knowledge. An empirical assessment of the results on the German credit dataset is also provided.