Towards a Reliable Evaluation of Local Interpretation Methods

Towards a Reliable Evaluation of Local Interpretation Methods
复制标题

对当地解释方法进行可靠的评估

DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
C. Ding
C. Ding
中科院分区:
--
文献类型:
--
作者:
Jun Li;Daoyu Lin;Yang Wang;Guangluan Xu;C. Ding

文献摘要

参考文献

被引文献

相似文献

深度神经网络在关键应用中的使用越来越多,使得可解释性迫切需要解决。局部解释方法是理解和解释深度神经网络的最普遍和最被接受的方法。如何有效地评价当地的解释方法是一个挑战。针对这一问题,提出了一个统一的评价框架,从准确性、说服性和类别区分性三个维度对局部解释方法进行评价。具体来说,为了评估正确性,我们设计了一个交互式用户特征注释工具,为本地解释方法提供地面实况。为了验证解释方法的有效性,我们迭代地显示部分解释结果,然后询问用户是否同意类别信息。同时,设计并构建了一套具有丰富层次结构的评价数据集。令人惊讶的是,一个发现是,现有的视觉解释方法不能满足所有的评价维度在同一时间,每一个都有自己的缺点。
The growing use of deep neural networks in critical applications is making interpretability urgently to be solved. Local interpretation methods are the most prevalent and accepted approach for understanding and interpreting deep neural networks. How to effectively evaluate the local interpretation methods is challenging. To address this question, a unified evaluation framework is proposed, which assesses local interpretation methods from three dimensions: accuracy, persuasibility and class discriminativeness. Specifically, in order to assess correctness, we designed an interactive user feature annotation tool to provide ground truth for local interpretation methods. To verify the usefulness of the interpretation method, we iteratively display part of the interpretation results, and then ask users whether they agree with the category information. At the same time, we designed and built a set of evaluation data sets with a rich hierarchical structure. Surprisingly, one finding is that the existing visual interpretation methods cannot satisfy all evaluation dimensions at the same time, and each has its own shortcomings.
DOI: 10.1073/pnas.1900654116
发表时间: 2019-10-29
影响因子: 11.1
作者:
Murdoch, W. James;Singh, Chandan;Yu, Bin
通讯作者: Yu, Bin