Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study.

Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs: A cross-sectional study.
复制标题

DOI:
10.1371/journal.pmed.1002683
复制
发表时间:
2018-11
期刊:
影响因子:
15.8
通讯作者:
Oermann EK
Oermann EK
中科院分区:
医学1区
文献类型:
--
作者:
Zech JR;Badgeley MA;Liu M;Costa AB;Titano JJ;Oermann EK

文献摘要

参考文献

被引文献

相似文献

人们对使用卷积神经网络(CNN)来分析医学成像以提供计算机辅助诊断(CAD)感兴趣。最近的工作表明,图像分类CNN可能无法像之前认为的那样推广到新数据。我们评估了CNN在模拟肺炎筛查任务中在三个医院系统中的推广情况。采用多个模型训练队列的横断面设计,使用分裂样本验证评价模型对外部研究中心的可推广性。共从三个机构抽取了158,323张胸片:美国国立卫生研究院临床中心(NIH; 30,805例患者中的112,120张)、西奈山医院(MSH; 12,904例患者中的42,396张)和印第安纳州大学患者护理网络(IU; 3,683例患者中的3,807张)。这些患者人群的平均年龄(SD)分别为46.9岁(16.6)、63.2岁(16.5)和49.6岁(17),女性百分比分别为43.5%、44.8%和57.3%。我们使用受试者工作特征曲线下面积(AUC)评估了与肺炎一致的放射学结果的个体模型,并使用DeLong检验比较了不同测试集的性能。相对于NIH和IU(1.2%和1.0%),MSH(34.2%)的肺炎患病率足够高,仅按医院系统分类即可在联合MSH-NIH数据集上实现AUC为0.861(95% CI 0.855-0.866)。在来自NIH或MSH的数据上训练的模型在IU上具有等同的性能(P值分别为0.580和0.273),并且相对于内部测试集(即,来自医院系统内的新数据用于训练数据; P值均<0.001)。通过结合来自MSH和NIH的训练和测试数据实现了最高的内部性能(AUC 0.931,95% CI 0.927-0.936),但该模型在IU时表现出显著较低的外部性能(AUC 0.815,95% CI 0.745-0.885,P = 0.001)。为了测试来自不同肺炎患病率的站点的汇总数据的效果,我们使用分层子抽样来生成MSH-NIH队列,这些队列仅在训练数据站点之间的疾病患病率方面存在差异。当两个训练数据中心具有相同的肺炎患病率时,该模型在外部IU数据上的表现一致(P = 0.88)。当研究中心之间的肺炎发生率存在10倍差异时,与平衡模型相比,内部测试性能有所改善(10× MSH风险P < 0.001; 10× NIH P = 0.002),但这种优越性未能推广到IU(MSH 10× P < 0.001; NIH 10× P = 0.027)。CNN能够直接检测99.95% NIH(22,050/22,062)和99.98% MSH(8,386/8,388)X光片的医院系统。我们的方法和可用的公共数据的主要限制是,我们不能完全评估其他因素可能导致医院系统特定的偏见。在5次自然比较中,有3次,筛选神经网络的内部性能优于外部性能。当模型在来自不同肺炎患病率的研究中心的汇总数据上进行训练时,它们在来自这些研究中心的新汇总数据上表现更好,但在外部数据上表现不佳。CNN可以强大地识别医院内的医院系统和部门,这些系统和部门在疾病负担方面可能存在很大差异,并可能混淆预测。Eric Oermann及其同事询问基于DL的肺炎检测模型是否在外部验证中表现良好,并考虑医院系统特定偏倚的影响。在X射线上使用卷积神经网络(CNN)诊断疾病的早期结果是有希望的,但尚未表明在一家医院或一组医院的X射线上训练的模型在不同的医院同样有效。在将这些工具用于现实世界临床环境中的计算机辅助诊断之前,我们必须验证它们在各种医院系统中的泛化能力。采用横断面设计对来自美国国立卫生研究院临床中心的158,323例胸部X射线进行肺炎筛查CNN的培训和评估(NIH; n = 112,120,来自30,805例患者)、西奈山医院(42,396,来自12,904例患者)和印第安纳州大学患者护理网络(n = 3,807,来自3,683例患者)。在3/5的自然比较中,来自外部医院的胸部X光片的性能显著低于来自原始医院系统的胸部X光片。CNN能够以极高的准确度检测到射线照片的获取位置(医院系统,医院部门),并相应地校准预测。CNN在X射线诊断疾病方面的表现不仅反映了它们在X射线上识别疾病特异性成像结果的能力,还反映了它们利用混淆信息的能力。基于用于模型训练的医院系统的测试数据对CNN性能的估计可能会夸大其可能的真实世界性能。
There is interest in using convolutional neural networks (CNNs) to analyze medical imaging to provide computer-aided diagnosis (CAD). Recent work has suggested that image classification CNNs may not generalize to new data as well as previously believed. We assessed how well CNNs generalized across three hospital systems for a simulated pneumonia screening task. A cross-sectional design with multiple model training cohorts was used to evaluate model generalizability to external sites using split-sample validation. A total of 158,323 chest radiographs were drawn from three institutions: National Institutes of Health Clinical Center (NIH; 112,120 from 30,805 patients), Mount Sinai Hospital (MSH; 42,396 from 12,904 patients), and Indiana University Network for Patient Care (IU; 3,807 from 3,683 patients). These patient populations had an age mean (SD) of 46.9 years (16.6), 63.2 years (16.5), and 49.6 years (17) with a female percentage of 43.5%, 44.8%, and 57.3%, respectively. We assessed individual models using the area under the receiver operating characteristic curve (AUC) for radiographic findings consistent with pneumonia and compared performance on different test sets with DeLong’s test. The prevalence of pneumonia was high enough at MSH (34.2%) relative to NIH and IU (1.2% and 1.0%) that merely sorting by hospital system achieved an AUC of 0.861 (95% CI 0.855–0.866) on the joint MSH–NIH dataset. Models trained on data from either NIH or MSH had equivalent performance on IU (P values 0.580 and 0.273, respectively) and inferior performance on data from each other relative to an internal test set (i.e., new data from within the hospital system used for training data; P values both <0.001). The highest internal performance was achieved by combining training and test data from MSH and NIH (AUC 0.931, 95% CI 0.927–0.936), but this model demonstrated significantly lower external performance at IU (AUC 0.815, 95% CI 0.745–0.885, P = 0.001). To test the effect of pooling data from sites with disparate pneumonia prevalence, we used stratified subsampling to generate MSH–NIH cohorts that only differed in disease prevalence between training data sites. When both training data sites had the same pneumonia prevalence, the model performed consistently on external IU data (P = 0.88). When a 10-fold difference in pneumonia rate was introduced between sites, internal test performance improved compared to the balanced model (10× MSH risk P < 0.001; 10× NIH P = 0.002), but this outperformance failed to generalize to IU (MSH 10× P < 0.001; NIH 10× P = 0.027). CNNs were able to directly detect hospital system of a radiograph for 99.95% NIH (22,050/22,062) and 99.98% MSH (8,386/8,388) radiographs. The primary limitation of our approach and the available public data is that we cannot fully assess what other factors might be contributing to hospital system–specific biases. Pneumonia-screening CNNs achieved better internal than external performance in 3 out of 5 natural comparisons. When models were trained on pooled data from sites with different pneumonia prevalence, they performed better on new pooled data from these sites but not on external data. CNNs robustly identified hospital system and department within a hospital, which can have large differences in disease burden and may confound predictions. Eric Oermann and colleagues ask whether a DL-based model for pneumonia detection performs well in external validation and consider the effects of hospital system–specific biases. Early results in using convolutional neural networks (CNNs) on X-rays to diagnose disease have been promising, but it has not yet been shown that models trained on X-rays from one hospital or one group of hospitals will work equally well at different hospitals. Before these tools are used for computer-aided diagnosis in real-world clinical settings, we must verify their ability to generalize across a variety of hospital systems. A cross-sectional design was used to train and evaluate pneumonia screening CNNs on 158,323 chest X-rays from the National Institutes of Health Clinical Center (NIH; n = 112,120 from 30,805 patients), Mount Sinai Hospital (42,396 from 12,904 patients), and Indiana University Network for Patient Care (n = 3,807 from 3,683 patients). In 3 out of 5 natural comparisons, performance on chest X-rays from outside hospitals was significantly lower than on held-out X-rays from the original hospital system. CNNs were able to detect where a radiograph was acquired (hospital system, hospital department) with extremely high accuracy and calibrate predictions accordingly. The performance of CNNs in diagnosing diseases on X-rays may reflect not only their ability to identify disease-specific imaging findings on X-rays but also their ability to exploit confounding information. Estimates of CNN performance based on test data from hospital systems used for model training may overstate their likely real-world performance.
DOI: 10.1148/radiol.2018171093
发表时间: 2018-05-01
期刊: RADIOLOGY
影响因子: 19.7
作者:
Zech, John;Pain, Margaret;Oermann, Eric Karl
通讯作者: Oermann, Eric Karl
DOI: 10.7326/m14-0698
发表时间: 2015-01-06
影响因子: 39.2
作者:
Collins, Gary S.;Reitsma, Johannes B.;Moons, Karel G. M.
通讯作者: Moons, Karel G. M.
DOI: 10.1136/bmj.j2835
发表时间: 2017-06-30
期刊: BMJ (Clinical research ed.)
影响因子: --
作者:
Pandis N;Chung B;Scherer RW;Elbourne D;Altman DG
通讯作者: Altman DG
DOI: 10.1001/jama.2017.18152
发表时间: 2017-12-12
影响因子: 120.7
作者:
Ting, Daniel Shu Wei;Cheung, Carol Yim-Lui;Wong, Tien Yin
通讯作者: Wong, Tien Yin
DOI: 10.1007/s10278-018-0053-3
发表时间: 2018-08-01
影响因子: 4.4
作者:
Mutasa, Simukayi;Chang, Peter D.;Ayyala, Rama
通讯作者: Ayyala, Rama