Improving Deep Learning Interpretability by Saliency Guided Training

Improving Deep Learning Interpretability by Saliency Guided Training
复制标题

DOI:
--
复制
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Ismail;H. C. Bravo;S. Feizi
A. Ismail;H. C. Bravo;S. Feizi
中科院分区:
其他
文献类型:
--
作者:
A. Ismail;H. C. Bravo;S. Feizi

文献摘要

被引文献

相似文献

显著性方法已被广泛用于突出模型预测中的重要输入特征。大多数现有的方法使用修改的梯度函数上的反向传播来生成显着图。因此,噪声梯度可能导致不忠实的特征属性。在本文中,我们解决了这个问题,并为神经网络引入了一个显着性指导训练过程,以减少预测中使用的噪声梯度,同时保留模型的预测性能。我们的显着性指导训练过程迭代地掩蔽具有小的和潜在的噪声梯度的特征,同时最大化掩蔽和未掩蔽输入的模型输出的相似性。我们将显着性引导的训练过程应用于来自计算机视觉,自然语言处理和不同神经架构(包括递归神经网络,卷积网络和变压器)的时间序列的各种合成和真实的数据集。通过定性和定量评估,我们表明显着性引导的训练过程显着提高了各个领域的模型可解释性,同时保持了其预测性能。
Saliency methods have been widely used to highlight important input features in model predictions. Most existing methods use backpropagation on a modified gradient function to generate saliency maps. Thus, noisy gradients can result in unfaithful feature attributions. In this paper, we tackle this issue and introduce a {\it saliency guided training}procedure for neural networks to reduce noisy gradients used in predictions while retaining the predictive performance of the model. Our saliency guided training procedure iteratively masks features with small and potentially noisy gradients while maximizing the similarity of model outputs for both masked and unmasked inputs. We apply the saliency guided training procedure to various synthetic and real data sets from computer vision, natural language processing, and time series across diverse neural architectures, including Recurrent Neural Networks, Convolutional Networks, and Transformers. Through qualitative and quantitative evaluations, we show that saliency guided training procedure significantly improves model interpretability across various domains while preserving its predictive performance.