Auditing the inference processes of medical-image classifiers by leveraging generative AI and the expertise of physicians.

Auditing the inference processes of medical-image classifiers by leveraging generative AI and the expertise of physicians.
复制标题

利用生成式人工智能和医生的专业知识来审核医学图像分类器的推理过程。

DOI:
10.1038/s41551-023-01160-9
复制
发表时间:
2023
影响因子:
28.1
通讯作者:
Lee,Su-In
Lee,Su-In
中科院分区:
工程技术1区
文献类型:
--
作者:
DeGrave,AlexJ;Cai,ZhuoRan;Janizek,JosephD;Daneshjou,Roxana;Lee,Su-In

文献摘要

相似文献

支持医学人工智能的大多数机器学习模型的推论很难解释。在这里,我们报告了一个模型审计的一般框架,它将医学专家的见解与一种高度可解释的人工智能的高度表达形式相结合。具体地说,我们利用皮肤科医生的专业知识,根据皮肤的皮肤镜和临床图像,区分黑色素瘤和黑色素瘤的“相貌相似”,以及生成模型呈现“反事实”图像的能力,以理解五种医学图像分类器的“推理”过程。通过改变图像属性来产生类似的图像,从而引发分类器的不同预测,并通过要求医生识别图像中具有医学意义的特征,反事实图像揭示了分类器既依赖于人类皮肤科医生使用的特征,如皮损色素沉着模式,也依赖于不受欢迎的特征,如背景皮肤纹理和颜色平衡。该框架可以应用于任何专业医学领域,使机器学习模型的强大推理过程在医学上是可理解的。
The inferences of most machine-learning models powering medical artificial intelligence are difficult to interpret. Here we report a general framework for model auditing that combines insights from medical experts with a highly expressive form of explainable artificial intelligence. Specifically, we leveraged the expertise of dermatologists for the clinical task of differentiating melanomas from melanoma ‘lookalikes’ on the basis of dermoscopic and clinical images of the skin, and the power of generative models to render ‘counterfactual’ images to understand the ‘reasoning’ processes of five medical-image classifiers. By altering image attributes to produce analogous images that elicit a different prediction by the classifiers, and by asking physicians to identify medically meaningful features in the images, the counterfactual images revealed that the classifiers rely both on features used by human dermatologists, such as lesional pigmentation patterns, and on undesirable features, such as background skin texture and colour balance. The framework can be applied to any specialized medical domain to make the powerful inference processes of machine-learning models medically understandable.