AI recognition of patient race in medical imaging: a modelling study.

AI recognition of patient race in medical imaging: a modelling study.
复制标题

DOI:
10.1016/s2589-7500(22)00063-2
复制
发表时间:
2022-06
影响因子:
30.8
通讯作者:
Zhang, Haoran
Zhang, Haoran
中科院分区:
医学1区
文献类型:
--
作者:
Gichoya, Judy Wawira;Banerjee, Imon;Bhimireddy, Ananth Reddy;Burns, John L.;Celi, Leo Anthony;Chen, Li-Ching;Correa, Ramon;Dullerud, Natalie;Ghassemi, Marzyeh;Huang, Shih-Cheng;Kuo, Po-Chih;Lungren, Matthew P.;Palmer, Lyle J.;Price, Brandon J.;Purkayastha, Saptarshi;Pyrros, Ayis T.;Oakden-Rayner, Lauren;Okechukwu, Chima;Seyyed-Kalantari, Laleh;Trivedi, Hari;Wang, Ryan;Zaiman, Zachary;Zhang, Haoran

文献摘要

被引文献

相似文献

以前在医学成像方面的研究表明,人工智能(AI)在检测一个人的种族方面具有不同的能力,但在医学成像上,人类专家在解释图像时,种族之间没有明显的相关性。我们的目标是对人工智能从医学图像中识别患者种族身份的能力进行全面评估。使用私人(埃默里CXR、埃默里胸部CT、埃默里颈椎和埃默里乳房X线检查)和公众(MIMIC-CXR,CheXpert,国家肺癌筛查试验,RSNA肺栓塞CT和数字手图谱)数据集,我们首先评估了深度学习模型在从医学图像中检测种族方面的性能量化,包括这些模型推广到外部环境和多种成像模式的能力。其次,我们评估了解剖学和表型人群特征的可能混淆,方法是评估这些假设的混淆因素使用回归模型孤立地检测种族的能力,并通过在由这些假设的混淆变量分层的数据集上测试深度学习模型来重新评估它们。最后,通过探索图像损坏对模型性能的影响,我们研究了AI模型识别种族的潜在机制。在我们的研究中,我们表明,标准的AI深度学习模型可以被训练来预测来自医学图像的种族,在多种成像模式下具有高性能,这在外部验证条件下(X射线成像[受试者工作特征曲线下的面积(AUC)范围为0.91 - 0.99],CT胸部成像[0.87 - 0.96]和乳房X光检查[0.81])。我们还表明,这种检测不是由于人种的代理或成像相关的替代协变量(例如,可能混杂因素的表现:体重指数[AUC 0·55],疾病分布[0·61]和乳腺密度[0·61])。最后,我们提供的证据表明,人工智能深度学习模型的能力在图像的所有解剖区域和频谱上持续存在,这表明在不受欢迎的情况下控制这种行为的努力将具有挑战性,需要进一步研究。我们的研究结果强调,人工智能深度学习模型预测自我报告种族的能力本身并不重要。然而,我们发现人工智能可以准确地预测自我报告的种族,即使是从损坏的,裁剪的和噪声的医学图像中,通常当临床专家不能时,也会为医学成像中的所有模型部署带来巨大的风险。国立生物医学影像与生物工程研究所、国立卫生研究院MIDRC基金、美国国家科学基金会、国立卫生研究院国家医学图书馆和台湾科技部。
Previous studies in medical imaging have shown disparate abilities of artificial intelligence (AI) to detect a person’s race, yet there is no known correlation for race on medical imaging that would be obvious to human experts when interpreting the images. We aimed to conduct a comprehensive evaluation of the ability of AI to recognise a patient’s racial identity from medical images. Using private (Emory CXR, Emory Chest CT, Emory Cervical Spine, and Emory Mammogram) and public (MIMIC-CXR, CheXpert, National Lung Cancer Screening Trial, RSNA Pulmonary Embolism CT, and Digital Hand Atlas) datasets, we evaluated, first, performance quantification of deep learning models in detecting race from medical images, including the ability of these models to generalise to external environments and across multiple imaging modalities. Second, we assessed possible confounding of anatomic and phenotypic population features by assessing the ability of these hypothesised confounders to detect race in isolation using regression models, and by re-evaluating the deep learning models by testing them on datasets stratified by these hypothesised confounding variables. Last, by exploring the effect of image corruptions on model performance, we investigated the underlying mechanism by which AI models can recognise race. In our study, we show that standard AI deep learning models can be trained to predict race from medical images with high performance across multiple imaging modalities, which was sustained under external validation conditions (x-ray imaging [area under the receiver operating characteristics curve (AUC) range 0·91–0·99], CT chest imaging [0·87–0·96], and mammography [0·81]). We also showed that this detection is not due to proxies or imaging-related surrogate covariates for race (eg, performance of possible confounders: body-mass index [AUC 0·55], disease distribution [0·61], and breast density [0·61]). Finally, we provide evidence to show that the ability of AI deep learning models persisted over all anatomical regions and frequency spectrums of the images, suggesting the efforts to control this behaviour when it is undesirable will be challenging and demand further study. The results from our study emphasise that the ability of AI deep learning models to predict self-reported race is itself not the issue of importance. However, our finding that AI can accurately predict self-reported race, even from corrupted, cropped, and noised medical images, often when clinical experts cannot, creates an enormous risk for all model deployments in medical imaging. National Institute of Biomedical Imaging and Bioengineering, MIDRC grant of National Institutes of Health, US National Science Foundation, National Library of Medicine of the National Institutes of Health, and Taiwan Ministry of Science and Technology.