Layer-Wise Relevance Propagation for Explaining Deep Neural Network Decisions in MRI-Based Alzheimer's Disease Classification

Layer-Wise Relevance Propagation for Explaining Deep Neural Network Decisions in MRI-Based Alzheimer's Disease Classification
复制标题

DOI:
10.3389/fnagi.2019.00194
复制
发表时间:
2019-07-31
影响因子:
4.8
通讯作者:
Ritter, Kerstin
Ritter, Kerstin
中科院分区:
医学2区
文献类型:
--
作者:
Boehle, Moritz;Eitel, Fabian;Ritter, Kerstin

文献摘要

被引文献

相似文献

深度神经网络在许多医学成像任务中产生了最先进的结果,包括基于结构磁共振成像(MRI)数据的阿尔茨海默病(AD)检测。然而,网络决策往往被认为是高度不透明的,这使得这些算法很难应用于临床常规。在这项研究中,我们建议使用分层相关传播(LRP)来可视化基于MRI数据的卷积神经网络对AD的决策。与其他可视化方法类似,LRP在输入空间中产生热图,指示每个体素对最终分类结果的重要性/相关性。与引导反向传播产生的敏感度图(“体素的哪个变化对结果影响最大?”)不同,LRP方法能够在输入空间直接突出对网络分类的积极贡献。特别是,我们证明了(1)LRP方法对于个体是非常特定的(“为什么这个人有AD?”)由于患者间的高度变异性,(2)在健康对照中与AD的相关性很小,(3)表现出很大相关性的区域与文献中已知的很好地相关。为了量化后者,我们计算每个大脑区域的总相关性的大小校正的度量,例如,相关性密度或相关性增益。虽然这些指标为AD患者提供了非常个性化的相关性模式的“指纹”,但非常重要的是包括海马体在内的颞叶区域。在讨论了一些限制,如对潜在模型和计算参数的敏感性后,我们得出结论,LRP可能具有很高的潜力来帮助临床医生解释基于结构性MRI数据诊断AD(以及潜在的其他疾病)的神经网络决策。
Deep neural networks have led to state-of-the-art results in many medical imaging tasks including Alzheimer's disease (AD) detection based on structural magnetic resonance imaging (MRI) data. However, the network decisions are often perceived as being highly non-transparent, making it difficult to apply these algorithms in clinical routine. In this study, we propose using layer-wise relevance propagation (LRP) to visualize convolutional neural network decisions for AD based on MRI data. Similarly to other visualization methods, LRP produces a heatmap in the input space indicating the importance/relevance of each voxel contributing to the final classification outcome. In contrast to susceptibility maps produced by guided backpropagation ("Which change in voxels would change the outcome most?"), the LRP method is able to directly highlight positive contributions to the network classification in the input space. In particular, we show that (1) the LRP method is very specific for individuals ("Why does this person have AD?") with high inter-patient variability, (2) there is very little relevance for AD in healthy controls and (3) areas that exhibit a lot of relevance correlate well with what is known from literature. To quantify the latter, we compute size-corrected metrics of the summed relevance per brain area, e.g., relevance density or relevance gain. Although these metrics produce very individual "fingerprints" of relevance patterns for AD patients, a lot of importance is put on areas in the temporal lobe including the hippocampus. After discussing several limitations such as sensitivity toward the underlying model and computation parameters, we conclude that LRP might have a high potential to assist clinicians in explaining neural network decisions for diagnosing AD (and potentially other diseases) based on structural MRI data.