SWAG-V: Explanations for Video using Superpixels Weighted by Average Gradients

SWAG-V: Explanations for Video using Superpixels Weighted by Average Gradients
复制标题

DOI:
10.1109/wacv51458.2022.00164
复制
发表时间:
2022-01
期刊:
2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Thomas Hartley;K. Sidorov;Christopher Willis;David Marshall
Thomas Hartley;K. Sidorov;Christopher Willis;David Marshall
中科院分区:
其他
文献类型:
--
作者:
Thomas Hartley;K. Sidorov;Christopher Willis;David Marshall

文献摘要

被引文献

相似文献

在开发解释技术时,将视频作为输入的CNN架构经常被忽视。尽管它们通常用于关键领域,如监视和医疗保健。为这些网络开发的解释技术必须考虑到额外的时间域,如果他们是成功的。在本文中,我们介绍了SWAG-V,这是SWAG的一个扩展,用于将视频作为输入的网络。此外,我们展示了这些解释如何可以创建这样一种方式,他们之间的平衡精细和粗糙的解释。通过创建包含输入视频帧的超像素,我们能够创建更好地定位对网络预测重要的输入区域的解释。我们比较SWAG-V对一些类似的技术使用的指标,如插入和删除,和弱本地化。我们使用Kinetics-400与C3 D和R(2+1)D网络架构计算这些,并发现SWAG-V能够胜过多种技术。
CNN architectures that take videos as an input are often overlooked when it comes to the development of explanation techniques. This is despite their use in often critical domains such as surveillance and healthcare. Explanation techniques developed for these networks must take into account the additional temporal domain if they are to be successful. In this paper we introduce SWAG-V, an extension of SWAG for use with networks that take video as an input. In addition we show how these explanations can be created in such a way that they are balanced between fine and coarse explanations. By creating superpixels that incorporate the frames of the input video we are able to create explanations that better locate regions of the input that are important to the networks prediction. We compare SWAG-V against a number of similar techniques using metrics such as insertion and deletion, and weak localisation. We compute these using Kinetics-400 with both the C3D and R(2+1)D network architectures and find that SWAG-V is able to outperform multiple techniques.