Interpretable Data-Based Explanations for Fairness Debugging

Interpretable Data-Based Explanations for Fairness Debugging
复制标题

DOI:
10.1145/3514221.3517886
复制
发表时间:
2021-12
期刊:
Proceedings of the 2022 International Conference on Management of Data
影响因子:
--
通讯作者:
Romila Pradhan;Jiongli Zhu;Boris Glavic;Babak Salimi
Romila Pradhan;Jiongli Zhu;Boris Glavic;Babak Salimi
中科院分区:
其他
文献类型:
--
作者:
Romila Pradhan;Jiongli Zhu;Boris Glavic;Babak Salimi

文献摘要

被引文献

相似文献

在文献中已经提出了各种各样的公平性度量和可解释的人工智能(XAI)方法,以识别在关键现实环境中使用的机器学习模型中的偏差。然而,仅仅报告模型的偏差或使用现有的XAI技术生成解释不足以定位并最终减轻偏差的来源。我们引入Gopher,一个系统,通过识别训练数据的连贯子集(这些子集是导致这种行为的根本原因),为偏差或意外的模型行为产生紧凑的、可解释的和因果的解释。具体来说,我们引入了因果责任的概念,它量化了通过删除或更新训练数据子集来干预训练数据可以解决偏差的程度。基于这一概念,我们开发了一种有效的方法来生成top-k模式,通过利用机器学习(ML)社区的技术来近似因果责任,并使用修剪规则来管理模式的大搜索空间,从而解释模型偏差。我们的实验评估表明,Gopher在产生可解释的解释,识别和调试源的偏见的有效性。
A wide variety of fairness metrics and eXplainable Artificial Intelligence (XAI) approaches have been proposed in the literature to identify bias in machine learning models that are used in critical real-life contexts. However, merely reporting on a model's bias or generating explanations using existing XAI techniques is insufficient to locate and eventually mitigate sources of bias. We introduce Gopher, a system that produces compact, interpretable, and causal explanations for bias or unexpected model behavior by identifying coherent subsets of the training data that are root-causes for this behavior. Specifically, we introduce the concept of causal responsibility that quantifies the extent to which intervening on training data by removing or updating subsets of it can resolve the bias. Building on this concept, we develop an efficient approach for generating the top-k patterns that explain model bias by utilizing techniques from the machine learning (ML) community to approximate causal responsibility, and using pruning rules to manage the large search space for patterns. Our experimental evaluation demonstrates the effectiveness of Gopher in generating interpretable explanations for identifying and debugging sources of bias.