Neural network attribution methods for problems in geoscience: A novel synthetic benchmark dataset

Neural network attribution methods for problems in geoscience: A novel synthetic benchmark dataset
复制标题

DOI:
10.1017/eds.2022.7
复制
发表时间:
2021-03
期刊:
Environmental Data Science
影响因子:
--
通讯作者:
Antonios Mamalakis;I. Ebert‐Uphoff;E. Barnes
Antonios Mamalakis;I. Ebert‐Uphoff;E. Barnes
中科院分区:
其他
文献类型:
--
作者:
Antonios Mamalakis;I. Ebert‐Uphoff;E. Barnes

文献摘要

被引文献

相似文献

摘要 尽管神经网络在地球科学的许多问题上应用日益成功,但其复杂的非线性结构使得对其预测的解释变得困难,这限制了模型的可信度,并且不允许科学家获得有关手头问题的物理洞察力。在可解释人工智能(XAI)这一新兴领域中已经引入了许多不同的方法,其目的是将网络的预测归因于输入域中的特定特征。XAI方法通常通过使用基准数据集(例如用于图像分类的MNIST或ImageNet)进行评估。然而,对于这些数据集的大多数,缺乏一种客观的、从理论推导得出的归因真值,这使得在许多情况下对XAI的评估具有主观性。此外,专门为地球科学问题设计的基准数据集也很罕见。在此,我们提供了一个基于可加性可分函数使用的框架,以便为回归问题生成归因基准数据集,对于这些问题,归因的真值是先验已知的。我们生成了一个大型基准数据集,并训练一个全连接网络来学习用于模拟的潜在函数。然后,我们将来自不同XAI方法的估计热图与真值进行比较,以确定特定XAI方法表现良好或不佳的实例。我们认为,像本文所介绍的这种归因基准对于神经网络在地球科学中的进一步应用,以及对于XAI方法更客观的评估和准确的实施都非常重要,这将提高模型的可信度并有助于发现新的科学知识。
Abstract Despite the increasingly successful application of neural networks to many problems in the geosciences, their complex and nonlinear structure makes the interpretation of their predictions difficult, which limits model trust and does not allow scientists to gain physical insights about the problem at hand. Many different methods have been introduced in the emerging field of eXplainable Artificial Intelligence (XAI), which aims at attributing the network’s prediction to specific features in the input domain. XAI methods are usually assessed by using benchmark datasets (such as MNIST or ImageNet for image classification). However, an objective, theoretically derived ground truth for the attribution is lacking for most of these datasets, making the assessment of XAI in many cases subjective. Also, benchmark datasets specifically designed for problems in geosciences are rare. Here, we provide a framework, based on the use of additively separable functions, to generate attribution benchmark datasets for regression problems for which the ground truth of the attribution is known a priori. We generate a large benchmark dataset and train a fully connected network to learn the underlying function that was used for simulation. We then compare estimated heatmaps from different XAI methods to the ground truth in order to identify examples where specific XAI methods perform well or poorly. We believe that attribution benchmarks as the ones introduced herein are of great importance for further application of neural networks in the geosciences, and for more objective assessment and accurate implementation of XAI methods, which will increase model trust and assist in discovering new science.