Revisiting the Trustworthiness of Saliency Methods in Radiology AI

Revisiting the Trustworthiness of Saliency Methods in Radiology AI
复制标题

DOI:
10.1148/ryai.220221
复制
发表时间:
2024-01-01
期刊:
RADIOLOGY-ARTIFICIAL INTELLIGENCE
影响因子:
--
通讯作者:
Yan, Pingkun
Yan, Pingkun
中科院分区:
其他
文献类型:
--
作者:
Zhang, Jiajin;Chao, Hanqing;Yan, Pingkun

文献摘要

被引文献

相似文献

目的:为了确定放射学人工智能(AI)中的显着性图是否容易受到输入的细微扰动的影响,这可能导致误导性解释,使用预测显着性相关性(PSC)来评估显着性方法的灵敏度和鲁棒性。材料与方法:在这项回顾性研究中,本地训练的深度学习模型和商业供应商提供的研究原型在来自CheXpert数据集的191229张胸片和来自人类脑肿瘤分类数据集的7022张MR图像上进行了系统评估。两名放射科医生对270对胸片进行了读片研究。一个模型不可知的方法计算PSC系数被用来评估七个常用的显着性方法的灵敏度和鲁棒性。结果如下:显著性方法在CheXpert数据集上具有低灵敏度(最大PSC,0.25; 95% CI:0.12,0.38)和弱稳健性(最大PSC,0.12; 95% CI:0.0,0.25),如通过利用本地训练的模型参数所证明的。进一步的评估表明,在不了解模型细节的情况下,从商业原型生成的显着性图可能与模型输出无关(在不影响显着性图的情况下,受试者工作特征曲线下的面积减少了8.6%)。人类观察者的研究证实,专家很难识别扰动图像;专家的正确率不到44.8%。结论:流行的显着性方法在两个扰动胸片数据集上的PSC值较低,表明灵敏度和鲁棒性较弱。所提出的PSC指标为验证医疗AI可解释性的可信度提供了一个有价值的量化工具。
Purpose: To determine whether saliency maps in radiology artificial intelligence (AI) are vulnerable to subtle perturbations of the input, which could lead to misleading interpretations, using prediction -saliency correlation (PSC) for evaluating the sensitivity and robustness of saliency methods. Materials and Methods: In this retrospective study, locally trained deep learning models and a research prototype provided by a commercial vendor were systematically evaluated on 191 229 chest radiographs from the CheXpert dataset and 7022 MR images from a human brain tumor classification dataset. Two radiologists performed a reader study on 270 chest radiograph pairs. A model -agnostic approach for computing the PSC coefficient was used to evaluate the sensitivity and robustness of seven commonly used saliency methods. Results: The saliency methods had low sensitivity (maximum PSC, 0.25; 95% CI: 0.12, 0.38) and weak robustness (maximum PSC, 0.12; 95% CI: 0.0, 0.25) on the CheXpert dataset, as demonstrated by leveraging locally trained model parameters. Further evaluation showed that the saliency maps generated from a commercial prototype could be irrelevant to the model output, without knowledge of the model specifics (area under the receiver operating characteristic curve decreased by 8.6% without affecting the saliency map). The human observer studies confirmed that it is difficult for experts to identify the perturbed images; the experts had less than 44.8% correctness. Conclusion: Popular saliency methods scored low PSC values on the two datasets of perturbed chest radiographs, indicating weak sensitivity and robustness. The proposed PSC metric provides a valuable quantification tool for validating the trustworthiness of medical AI explainability.