SWAG: Superpixels Weighted by Average Gradients for Explanations of CNNs

SWAG: Superpixels Weighted by Average Gradients for Explanations of CNNs
复制标题

DOI:
10.1109/wacv48630.2021.00047
复制
发表时间:
2021-01
期刊:
2021 IEEE Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Thomas Hartley;K. Sidorov;C. Willis;David Marshall
Thomas Hartley;K. Sidorov;C. Willis;David Marshall
中科院分区:
其他
文献类型:
--
作者:
Thomas Hartley;K. Sidorov;C. Willis;David Marshall

文献摘要

相似文献

在医学图像分析、监控和自动驾驶等领域,提供准确且可解释的 CNN 运行解释变得至关重要。在这些领域,重要的是要确信 CNN 正在按预期工作,并且显着图的解释提供了一种有效的方法。在本文中,我们提出了一对互补的贡献,这些贡献在准确性和实用性方面改进了基于区域的解释的现有技术。第一个是 SWAG,这是一种使用超像素对判别区域快速生成准确解释的方法,这对于 Grad-CAM、LIME 或其他基于区域的方法来说是一种更准确、更高效、更可调的替代方法。第二个贡献基于对如何最好地生成用于表示图像中发现的特征的超像素的研究。使用 SWAG,我们使用从图像创建的超像素、图像和反向传播梯度的组合以及梯度本身进行比较。据我们所知,这是第一个提出的方法,使用显式创建的超像素来生成解释,以表示对网络重要的判别特征。为了进行比较,我们使用 ImageNet 和具有挑战性的细粒度数据集来衡量一系列指标。我们通过实验证明,与 Grad-CAM、Grad-CAM++、LIME、XRAI 和 RISE 相比,我们的方法提供了最佳的局部和全局精度。
Providing an explanation of the operation of CNNs that is both accurate and interpretable is becoming essential in fields like medical image analysis, surveillance, and autonomous driving. In these areas, it is important to have confidence that the CNN is working as expected and explanations from saliency maps provide an efficient way of doing this. In this paper, we propose a pair of complementary contributions that improve upon the state of the art for region-based explanations in both accuracy and utility. The first is SWAG, a method for generating accurate explanations quickly using superpixels for discriminative regions which is meant to be a more accurate, efficient, and tunable drop in replacement method for Grad-CAM, LIME, or other region-based methods. The second contribution is based on an investigation into how to best generate the superpixels used to represent the features found within the image. Using SWAG, we compare using superpixels created from the image, a combination of the image and backpropagated gradients, and the gradients themselves. To the best of our knowledge, this is the first method proposed to generate explanations using superpixels explicitly created to represent the discriminative features important to the network. To compare we use both ImageNet and challenging fine-grained datasets over a range of metrics. We demonstrate experimentally that our methods provide the best local and global accuracy compared to Grad-CAM, Grad-CAM++, LIME, XRAI, and RISE.