Framework for Evaluating Faithfulness of Local Explanations

Framework for Evaluating Faithfulness of Local Explanations
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
S. Dasgupta;Nave Frost;Michal Moshkovitz
S. Dasgupta;Nave Frost;Michal Moshkovitz
中科院分区:
其他
文献类型:
--
作者:
S. Dasgupta;Nave Frost;Michal Moshkovitz

文献摘要

相似文献

我们研究解释系统对底层预测模型的忠实度。我们表明,这可以捕获的两个性质,一致性和充分性,并引入定量措施的程度,这些持有。有趣的是,这些度量依赖于测试时数据分布。对于各种现有的解释系统,如锚,我们分析研究这些量。我们还提供了估计量和样本复杂度界,以经验确定黑箱解释系统的可靠性。最后,我们用实验验证了新的性质和估计量。
We study the faithfulness of an explanation system to the underlying prediction model. We show that this can be captured by two properties, consistency and sufficiency, and introduce quantitative measures of the extent to which these hold. Interestingly, these measures depend on the test-time data distribution. For a variety of existing explanation systems, such as anchors, we analytically study these quantities. We also provide estimators and sample complexity bounds for empirically determining the faithfulness of black-box explanation systems. Finally, we experimentally validate the new properties and estimators.