Improving performance of deep learning models with axiomatic attribution priors and expected gradients

Improving performance of deep learning models with axiomatic attribution priors and expected gradients
复制标题

DOI:
10.1038/s42256-021-00343-w
复制
发表时间:
2021-05-31
影响因子:
23.8
通讯作者:
Lee, Su-In
Lee, Su-In
中科院分区:
计算机科学1区
文献类型:
--
作者:
Erion, Gabriel;Janizek, Joseph D.;Lee, Su-In

文献摘要

被引文献

相似文献

神经网络在各个领域的应用越来越受欢迎,但在实践中,需要进一步的方法来确保模型是与该领域的先验知识一致的学习模式。一种新的方法引入了一种称为“期望梯度”的解释方法,该方法可以使用理论驱动的特征归因先验进行训练,以提高模型在现实世界任务中的性能。最近的研究表明,深度网络的特征归因方法本身可以纳入训练;这些归因先验对具有某些理想属性的模型进行优化——最常见的是,特定的特征是重要的或不重要的。这些归因先验通常基于不能保证满足理想的可解释性公理的归因方法,例如完整性和实现不变性。在这里,我们引入归因先验来优化解释的更高级别属性,如平滑性和稀疏性,这是由一种快速的新的归因方法公式实现的,称为期望梯度,它满足许多重要的可解释性公理。这提高了模型在许多现实世界任务中的性能,在这些任务中,之前的归因先验是失败的。我们的实验表明,将高水平归因先验与预期梯度归因相结合的结果在图像、基因表达和医疗保健数据集上是一致的。我们相信,这项工作激励并提供了必要的工具,以支持在应用机器学习的许多领域广泛采用公理归因先验。实现和我们的结果已经免费提供给学术界。
Neural networks are becoming increasingly popular for applications in various domains, but in practice, further methods are necessary to make sure the models are learning patterns that agree with prior knowledge about the domain. A new approach introduces an explanation method, called 'expected gradients', that enables training with theoretically motivated feature attribution priors, to improve model performance on real-world tasks.Recent research has demonstrated that feature attribution methods for deep networks can themselves be incorporated into training; these attribution priors optimize for a model whose attributions have certain desirable properties-most frequently, that particular features are important or unimportant. These attribution priors are often based on attribution methods that are not guaranteed to satisfy desirable interpretability axioms, such as completeness and implementation invariance. Here we introduce attribution priors to optimize for higher-level properties of explanations, such as smoothness and sparsity, enabled by a fast new attribution method formulation called expected gradients that satisfies many important interpretability axioms. This improves model performance on many real-world tasks where previous attribution priors fail. Our experiments show that the gains from combining higher-level attribution priors with expected gradients attributions are consistent across image, gene expression and healthcare datasets. We believe that this work motivates and provides the necessary tools to support the widespread adoption of axiomatic attribution priors in many areas of applied machine learning. The implementations and our results have been made freely available to academic communities.