Concept whitening for interpretable image recognition

Concept whitening for interpretable image recognition
复制标题

DOI:
10.1038/s42256-020-00265-z
复制
发表时间:
2020-12-01
影响因子:
23.8
通讯作者:
Rudin, Cynthia
Rudin, Cynthia
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chen, Zhi;Bei, Yijie;Rudin, Cynthia

文献摘要

被引文献

相似文献

当我们遍历这些层时,神经网络对一个概念进行了什么编码?机器学习中的可解释性无疑是重要的,但神经网络的计算非常难以理解。试图看到其隐藏层的内部可能会产生误导,无法使用或依赖于潜在空间拥有它可能不具有的属性。在这里,我们不是试图事后分析神经网络,而是引入一种称为概念白化(CW)的机制来改变网络的给定层,以使我们更好地理解导致该层的计算。当将概念白化模块添加到卷积神经网络中时,潜在空间被白化(即去相关和归一化),并且潜在空间的轴线与已知感兴趣的概念对齐。通过实验,我们表明CW可以让我们更清楚地理解网络是如何在层上逐渐学习概念的。CW是批归一化层的替代方案,因为它对潜在空间进行归一化和去相关(漂白)。CW可以用于网络的任何层,而不会影响预测性能。人们对“可解释的”人工智能很感兴趣,但大多数努力都是关于事后方法的。相反,神经网络可以通过一种方法使人类可以理解的概念(飞机、床、灯等)沿着其潜在空间的轴线排列,从而具有内在的可解释性。
What does a neural network encode about a concept as we traverse through the layers? Interpretability in machine learning is undoubtedly important, but the calculations of neural networks are very challenging to understand. Attempts to see inside their hidden layers can be misleading, unusable or rely on the latent space to possess properties that it may not have. Here, rather than attempting to analyse a neural network post hoc, we introduce a mechanism, called concept whitening (CW), to alter a given layer of the network to allow us to better understand the computation leading up to that layer. When a concept whitening module is added to a convolutional neural network, the latent space is whitened (that is, decorrelated and normalized) and the axes of the latent space are aligned with known concepts of interest. By experiment, we show that CW can provide us with a much clearer understanding of how the network gradually learns concepts over layers. CW is an alternative to a batch normalization layer in that it normalizes, and also decorrelates (whitens), the latent space. CW can be used in any layer of the network without hurting predictive performance.There is much interest in 'explainable' AI, but most efforts concern post hoc methods. Instead, a neural network can be made inherently interpretable, with an approach that involves making human-understandable concepts (aeroplane, bed, lamp and so on) align along the axes of its latent space.