Fairness via Explanation Quality: Evaluating Disparities in the Quality of Post hoc Explanations

Fairness via Explanation Quality: Evaluating Disparities in the Quality of Post hoc Explanations
复制标题

DOI:
10.1145/3514094.3534159
复制
发表时间:
2022-05
期刊:
Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society
影响因子:
--
通讯作者:
Jessica Dai;Sohini Upadhyay;U. Aïvodji;Stephen H. Bach;Himabindu Lakkaraju
Jessica Dai;Sohini Upadhyay;U. Aïvodji;Stephen H. Bach;Himabindu Lakkaraju
中科院分区:
其他
文献类型:
--
作者:
Jessica Dai;Sohini Upadhyay;U. Aïvodji;Stephen H. Bach;Himabindu Lakkaraju

文献摘要

相似文献

由于事后解释方法越来越多地被用来解释高风险环境中的复杂模型,因此确保所得到的解释的质量在人群的所有子群体中始终保持高质量变得至关重要。例如,与属于例如,女性,比其他性别的人更不准确。在这项工作中,我们发起了基于群体的解释质量差异的研究。为此,我们首先概述了几个关键属性,有助于解释质量,即保真度(准确性),稳定性,一致性和稀疏性,并讨论为什么以及如何在这些属性的差异可能是特别有问题的。然后,我们提出了一个评估框架,可以定量测量的解释质量的差异。使用这个框架,我们进行了实证分析,三个数据集,六个事后解释方法,和不同的模型类,以了解是否以及何时出现基于群体的解释质量的差异。我们的研究结果表明,当被解释的模型是复杂的和非线性的时,这种差异更有可能发生。我们还观察到,某些事后解释方法(例如,整合型人格(SHAP)更有可能表现出差异。我们的工作揭示了以前未探索的方式,解释方法可能会引入不公平的真实的世界的决策。
As post hoc explanation methods are increasingly being leveraged to explain complex models in high-stakes settings, it becomes critical to ensure that the quality of the resulting explanations is consistently high across all subgroups of a population. For instance, it should not be the case that explanations associated with instances belonging to, e.g., women, are less accurate than those associated with other genders. In this work, we initiate the study of identifying group-based disparities in explanation quality. To this end, we first outline several key properties that contribute to explanation quality-namely, fidelity (accuracy), stability, consistency, and sparsity-and discuss why and how disparities in these properties can be particularly problematic. We then propose an evaluation framework which can quantitatively measure disparities in the quality of explanations. Using this framework, we carry out an empirical analysis with three datasets, six post hoc explanation methods, and different model classes to understand if and when group-based disparities in explanation quality arise. Our results indicate that such disparities are more likely to occur when the models being explained are complex and non-linear. We also observe that certain post hoc explanation methods (e.g., Integrated Gradients, SHAP) are more likely to exhibit disparities. Our work sheds light on previously unexplored ways in which explanation methods may introduce unfairness in real world decision making.