Reliable Post hoc Explanations: Modeling Uncertainty in Explainability

Reliable Post hoc Explanations: Modeling Uncertainty in Explainability
复制标题

DOI:
--
复制
发表时间:
2020-08
期刊:
--
影响因子:
--
通讯作者:
Dylan Slack;Sophie Hilgard;Sameer Singh;Himabindu Lakkaraju
Dylan Slack;Sophie Hilgard;Sameer Singh;Himabindu Lakkaraju
中科院分区:
其他
文献类型:
--
作者:
Dylan Slack;Sophie Hilgard;Sameer Singh;Himabindu Lakkaraju

文献摘要

相似文献

随着黑箱解释越来越多地被用于在高风险环境中建立模型可信度,确保这些解释准确可靠非常重要。然而,先前的工作表明,最先进的技术产生的解释是不一致的,不稳定的,并提供很少的洞察其正确性和可靠性。此外,这些方法在计算上也是低效的,并且需要显著的超参数调整。在本文中,我们通过开发一种新的贝叶斯框架来解决上述挑战,该框架用于生成本地解释沿着其相关的不确定性。我们实例化这个框架,以获得贝叶斯版本的LIME和KernelSHAP输出可信区间的功能的重要性,捕捉相关的不确定性。由此产生的解释不仅使我们能够对它们的质量做出具体的推断(例如,有95%的机会特征重要性位于给定范围内),但也是高度一致和稳定的。我们进行了详细的理论分析,利用上述不确定性来估计要采样多少扰动,以及如何采样以加快收敛。这项工作首次尝试用流行的解释方法一次性解决几个关键问题,从而以计算效率高的方式产生一致,稳定和可靠的解释。多个真实的世界数据集和用户研究的实验评估表明,所提出的框架的有效性。
As black box explanations are increasingly being employed to establish model credibility in high-stakes settings, it is important to ensure that these explanations are accurate and reliable. However, prior work demonstrates that explanations generated by state-of-the-art techniques are inconsistent, unstable, and provide very little insight into their correctness and reliability. In addition, these methods are also computationally inefficient, and require significant hyper-parameter tuning. In this paper, we address the aforementioned challenges by developing a novel Bayesian framework for generating local explanations along with their associated uncertainty. We instantiate this framework to obtain Bayesian versions of LIME and KernelSHAP which output credible intervals for the feature importances, capturing the associated uncertainty. The resulting explanations not only enable us to make concrete inferences about their quality (e.g., there is a 95% chance that the feature importance lies within the given range), but are also highly consistent and stable. We carry out a detailed theoretical analysis that leverages the aforementioned uncertainty to estimate how many perturbations to sample, and how to sample for faster convergence. This work makes the first attempt at addressing several critical issues with popular explanation methods in one shot, thereby generating consistent, stable, and reliable explanations with guarantees in a computationally efficient manner. Experimental evaluation with multiple real world datasets and user studies demonstrate that the efficacy of the proposed framework.