Identifying Interpretable Subspaces in Image Representations

Identifying Interpretable Subspaces in Image Representations
复制标题

DOI:
10.48550/arxiv.2307.10504
复制
发表时间:
2023-07
期刊:
ArXiv
影响因子:
--
通讯作者:
N. Kalibhat;S. Bhardwaj;Bayan Bruss;Hamed Firooz;Maziar Sanjabi;S. Feizi
N. Kalibhat;S. Bhardwaj;Bayan Bruss;Hamed Firooz;Maziar Sanjabi;S. Feizi
中科院分区:
其他
文献类型:
--
作者:
N. Kalibhat;S. Bhardwaj;Bayan Bruss;Hamed Firooz;Maziar Sanjabi;S. Feizi

文献摘要

相似文献

我们提出了使用对比概念的自动特征解释(Automatic Feature Explanation using Contrasting Concepts,简称CAPCON),这是一个解释图像表征特征的可解释性框架。对于目标功能,CSPCON使用大型字幕数据集(如LAION-400 m)和预训练的视觉语言模型(如CLIP)为其高度活跃的裁剪图像添加字幕。标题中的每个单词都被评分和排名,从而产生少量共享的、人类可理解的概念,这些概念紧密地描述了目标特征。ESTCON还使用低激活(反事实)图像进行对比解释,以消除虚假概念。虽然许多现有的方法独立地解释特征,但我们在最先进的自监督和监督模型中观察到,只有不到20%的表示空间可以由单个特征解释。我们发现,在较大的空间中的功能变得更可解释的组中进行研究时,可以解释高阶评分的概念,通过ESTCON。我们将讨论如何提取的概念可以用来解释和调试下游任务中的失败。最后,我们提出了一种技术,通过学习一个简单的线性变换,将概念从一个(可解释的)表示空间转移到另一个看不见的表示空间。代码可在https://github.com/NehaKalibhat/falcon-explain上获得。
We propose Automatic Feature Explanation using Contrasting Concepts (FALCON), an interpretability framework to explain features of image representations. For a target feature, FALCON captions its highly activating cropped images using a large captioning dataset (like LAION-400m) and a pre-trained vision-language model like CLIP. Each word among the captions is scored and ranked leading to a small number of shared, human-understandable concepts that closely describe the target feature. FALCON also applies contrastive interpretation using lowly activating (counterfactual) images, to eliminate spurious concepts. Although many existing approaches interpret features independently, we observe in state-of-the-art self-supervised and supervised models, that less than 20% of the representation space can be explained by individual features. We show that features in larger spaces become more interpretable when studied in groups and can be explained with high-order scoring concepts through FALCON. We discuss how extracted concepts can be used to explain and debug failures in downstream tasks. Finally, we present a technique to transfer concepts from one (explainable) representation space to another unseen representation space by learning a simple linear transformation. Code available at https://github.com/NehaKalibhat/falcon-explain.