Capuchin: Causal Database Repair for Algorithmic Fairness

Capuchin: Causal Database Repair for Algorithmic Fairness
复制标题

Capuchin:因果数据库修复以实现算法公平

DOI:
--
复制
发表时间:
2019
期刊:
arXiv.org
影响因子:
--
通讯作者:
Dan Suciu
Dan Suciu
中科院分区:
--
文献类型:
--
作者:
Babak Salimi;Luke Rodriguez;Bill Howe;Dan Suciu

文献摘要

参考文献

被引文献

相似文献

公平性越来越被认为是机器学习系统的关键组成部分。然而,训练这些系统所依据的基础数据往往反映出歧视,表明存在数据库修复问题。现有的公平处理方法依赖于统计相关性,而统计相关性可能会被辛普森悖论等统计异常所愚弄。基于因果关系的公平性定义的提案可以正确地对其中一些情况进行建模,但它们需要对底层因果模型进行规范。在本文中,我们将这种情况形式化为数据库修复问题,证明了公平分类器在可接受变量方面的充分条件,而不是完整的因果模型。我们证明这些条件正确地捕获了微妙的公平违规行为。然后,我们使用这些条件作为数据库修复算法的基础,该算法为在其训练标签上训练的分类器提供可证明的公平性保证。我们在真实数据上评估我们的算法,证明在文献中提出的多种公平性指标上比现有技术有所改进,同时保持高实用性。
Fairness is increasingly recognized as a critical component of machine learning systems. However, it is the underlying data on which these systems are trained that often reflect discrimination, suggesting a database repair problem. Existing treatments of fairness rely on statistical correlations that can be fooled by statistical anomalies, such as Simpson's paradox. Proposals for causality-based definitions of fairness can correctly model some of these situations, but they require specification of the underlying causal models. In this paper, we formalize the situation as a database repair problem, proving sufficient conditions for fair classifiers in terms of admissible variables as opposed to a complete causal model. We show that these conditions correctly capture subtle fairness violations. We then use these conditions as the basis for database repair algorithms that provide provable fairness guarantees about classifiers trained on their training labels. We evaluate our algorithms on real data, demonstrating improvement over the state of the art on multiple fairness metrics proposed in the literature while retaining high utility.
大数据警务的不同影响
DOI: --
发表时间: 2017
期刊: Georgia law review
影响因子: --
作者:
Selbst, Andrew D.
通讯作者: Selbst, Andrew D.