Concept Gradient: Concept-based Interpretation Without Linear Assumption

Concept Gradient: Concept-based Interpretation Without Linear Assumption
复制标题

DOI:
10.48550/arxiv.2208.14966
复制
发表时间:
2022-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Andrew Bai;Chih-Kuan Yeh;Pradeep Ravikumar;Neil Y. C. Lin;Cho-Jui Hsieh
Andrew Bai;Chih-Kuan Yeh;Pradeep Ravikumar;Neil Y. C. Lin;Cho-Jui Hsieh
中科院分区:
其他
文献类型:
--
作者:
Andrew Bai;Chih-Kuan Yeh;Pradeep Ravikumar;Neil Y. C. Lin;Cho-Jui Hsieh

文献摘要

被引文献

相似文献

对黑盒模型的基于概念的解释通常对人类来说更直观。概念激活向量(CAV)是最常用的概念解释方法。CAV依赖于学习给定模型和概念的某些潜在表示之间的线性关系。线性可分性通常是隐含假设的,但一般不成立。在这项工作中,我们从基于概念的解释的初衷,并提出了概念梯度(CG),扩展了基于概念的解释线性概念函数。我们表明,对于一个一般的(潜在的非线性)概念,我们可以在数学上评估一个小的概念变化如何影响模型的预测,这导致了基于梯度的解释扩展到概念空间。我们的经验表明,CG优于CAV在玩具的例子和真实的世界的数据集。
Concept-based interpretations of black-box models are often more intuitive for humans to understand. The most widely adopted approach for concept-based interpretation is Concept Activation Vector (CAV). CAV relies on learning a linear relation between some latent representation of a given model and concepts. The linear separability is usually implicitly assumed but does not hold true in general. In this work, we started from the original intent of concept-based interpretation and proposed Concept Gradient (CG), extending concept-based interpretation beyond linear concept functions. We showed that for a general (potentially non-linear) concept, we can mathematically evaluate how a small change of concept affecting the model's prediction, which leads to an extension of gradient-based interpretation to the concept space. We demonstrated empirically that CG outperforms CAV in both toy examples and real world datasets.