Failure Detection in Deep Neural Networks for Medical Imaging.

Failure Detection in Deep Neural Networks for Medical Imaging.
复制标题

用于医学成像的深度神经网络故障检测。

DOI:
10.3389/fmedt.2022.919046
复制
发表时间:
2022
影响因子:
--
通讯作者:
Rasool, Ghulam
Rasool, Ghulam
中科院分区:
其他
文献类型:
--
作者:
Ahmed, Sabeen;Dera, Dimah;Hassan, Saud Ul;Bouaynaya, Nidhal;Rasool, Ghulam

文献摘要

被引文献

相似文献

深度神经网络(DNN)已经开始在现代医疗系统中发挥作用。DNN正在开发用于各种疾病的诊断、预后、治疗计划和结果预测。随着DNN在现代医疗保健中的应用越来越多,其可信度和可靠性变得越来越重要。可信度的一个重要方面是检测医疗环境中部署的DNN的性能下降和故障。DNN产生的softmax输出值不是模型置信度的校准度量。Softmax概率数字通常高于实际模型置信度。模型置信度-准确度差距因错误预测和噪声输入而进一步增加。我们使用最近提出的贝叶斯深度神经网络(BDNN)来学习模型参数中的不确定性。这些模型同时输出预测和预测的置信度。通过在各种噪声条件下测试这些模型,我们证明了(学习的)预测置信度得到了很好的校准。我们使用这些可靠的置信度值来监控DNN中的性能下降和故障检测。我们提出了两种不同的故障检测方法。在第一种方法中,我们定义了一个固定的阈值的基础上的行为的预测置信度与变化的信噪比(SNR)的测试数据集。第二种方法利用神经网络学习阈值。所提出的故障检测机制无缝地放弃决策时,BDNN的信心是低于定义的阈值,并举行人工审查的决定。结果,模型的准确性提高了看不见的测试样本。我们在三个医学成像数据集上测试了我们提出的方法:PathMNIST,DermaMNIST和OrganAMNIST,在不同的噪声水平和类型下。测试图像的噪声的增加增加了弃权样本的数量。BDNN具有固有的鲁棒性,并且使用所提出的故障检测方法显示出超过10%的准确性提高。弃权样本数量的增加或预测方差的突然增加表明模型性能下降或可能失败。我们的工作有可能提高DNN的可信度,并增强用户对模型预测的信心。
Deep neural networks (DNNs) have started to find their role in the modern healthcare system. DNNs are being developed for diagnosis, prognosis, treatment planning, and outcome prediction for various diseases. With the increasing number of applications of DNNs in modern healthcare, their trustworthiness and reliability are becoming increasingly important. An essential aspect of trustworthiness is detecting the performance degradation and failure of deployed DNNs in medical settings. The softmax output values produced by DNNs are not a calibrated measure of model confidence. Softmax probability numbers are generally higher than the actual model confidence. The model confidence-accuracy gap further increases for wrong predictions and noisy inputs. We employ recently proposed Bayesian deep neural networks (BDNNs) to learn uncertainty in the model parameters. These models simultaneously output the predictions and a measure of confidence in the predictions. By testing these models under various noisy conditions, we show that the (learned) predictive confidence is well calibrated. We use these reliable confidence values for monitoring performance degradation and failure detection in DNNs. We propose two different failure detection methods. In the first method, we define a fixed threshold value based on the behavior of the predictive confidence with changing signal-to-noise ratio (SNR) of the test dataset. The second method learns the threshold value with a neural network. The proposed failure detection mechanisms seamlessly abstain from making decisions when the confidence of the BDNN is below the defined threshold and hold the decision for manual review. Resultantly, the accuracy of the models improves on the unseen test samples. We tested our proposed approach on three medical imaging datasets: PathMNIST, DermaMNIST, and OrganAMNIST, under different levels and types of noise. An increase in the noise of the test images increases the number of abstained samples. BDNNs are inherently robust and show more than 10% accuracy improvement with the proposed failure detection methods. The increased number of abstained samples or an abrupt increase in the predictive variance indicates model performance degradation or possible failure. Our work has the potential to improve the trustworthiness of DNNs and enhance user confidence in the model predictions.