Deep Learning for Chest Radiograph Diagnosis in the Emergency Department

Deep Learning for Chest Radiograph Diagnosis in the Emergency Department
复制标题

DOI:
10.1148/radiol.2019191225
复制
发表时间:
2019-12-01
期刊:
影响因子:
19.7
通讯作者:
Park, Chang Min
Park, Chang Min
中科院分区:
医学1区
文献类型:
--
作者:
Hwang, Eui Jin;Nam, Ju Gang;Park, Chang Min

文献摘要

被引文献

相似文献

背景资料:深度学习(DL)算法的性能应在其临床实施之前在实际临床情况中进行验证。目的:评估DL算法在急诊科(艾德)环境中识别具有临床相关异常的胸片的性能。材料和方法:这首单曲-中心回顾性研究包括在1月1日至3月31日期间访问艾德并接受初始胸部X线摄影的连续患者,2017.采用市售的DL算法分析胸片,通过确定受试者工作特征曲线下面积(AUC)、灵敏度和特异性(在预定义的工作临界值(高灵敏度和高特异性临界值))来评价算法的性能。该算法的灵敏度和特异性进行了比较,与那些随叫随到的放射科居民解释胸片在实际工作中使用McNemar测试。如果有不一致的结果之间的算法和居民,居民重新解释的胸片通过使用算法的output.Results:共1135例患者(平均年龄,53岁+/- 18; 582名男性)进行了评价。在识别异常胸片时,算法显示AUC为0.95(95%置信区间[CI]:0.93,0.96),灵敏度为88.7%(256张X线片中的227张; 95% CI:84.1%,92.3%),特异性为69.6%(612/879张X线片; 95% CI:66.5%,72.7%),高灵敏度临界值和灵敏度81.6%(209/256张X线片; 95% CI:76.3%,86.2%),在高特异性临界值时的特异性为90.3%(794/879张X线片; 95% CI:88.2%,92.2%)。与算法相比,放射科住院医师显示出较低的灵敏度(65.6% [256张X光片中的168张; 95% CI:59.5%,71.4%],P < .001)和较高的特异性(98.1% [879张X光片中的862张; 95% CI:96.9%,98.9%],P < .001)。在使用算法的输出重新解释胸部X光片后,居民的敏感性提高了(73.4% [188/256; 95%CI:68.0%,78.8%],P = .003),而特异性降低(94.3% [879张X光片中的829张; 95% CI:92.8%,95.8%],P < .001)。结论:急诊科胸部X光片使用的深度学习算法显示了识别临床相关异常的诊断性能,并有助于提高放射科住院医师评估的灵敏度。在CC BY 4.0许可证下发布。
Background: The performance of a deep learning (DL) algorithm should be validated in actual clinical situations, before its clinical implementation.Purpose: To evaluate the performance of a DL algorithm for identifying chest radiographs with clinically relevant abnormalities in the emergency department (ED) setting.Materials and Methods: This single-center retrospective study included consecutive patients who visited the ED and underwent initial chest radiography between January 1 and March 31, 2017. Chest radiographs were analyzed with a commercially available DL algorithm.The performance of the algorithm was evaluated by determining the area under the receiver operating characteristic curve (AUC), sensitivity, and specificity at predefined operating cutoffs (high-sensitivity and high-specificity cutoffs). The sensitivities and specificities of the algorithm were compared with those of the on-call radiology residents who interpreted the chest radiographs in the actual practice by using McNemar tests. If there were discordant findings between the algorithm and resident, the residents reinterpreted the chest radiographs by using the algorithm's output.Results: A total of 1135 patients (mean age, 53 years +/- 18; 582 men) were evaluated. In the identification of abnormal chest radiographs,the algorithm showed an AUC of 0.95 (95% confidence interval [CI]: 0.93, 0.96), a sensitivity of 88.7% (227 of 256 radiographs; 95% CI: 84.1%, 92.3%), and a specificity of 69.6% (612 of 879 radiographs; 95% CI: 66.5%, 72.7%) at the high sensitivity cutoff and a sensitivity of 81.6% (209 of 256 radiographs; 95% CI: 76.3%, 86.2%) and specificity of 90.3% (794 of 879 radiographs; 95% CI: 88.2%, 92.2%) at the high-specificity cutoff. Radiology residents showed lower sensitivity (65.6% [168 of 256 radiographs; 95% CI: 59.5%, 71.4%], P < .001) and higher specificity (98.1% [862 of 879 radiographs; 95% CI: 96.9%,98.9%], P < .001) compared with the algorithm. After reinterpretation of chest radiographs with use of the algorithm's outputs,the sensitivity of the residents improved (73.4% [188 of 256 radiographs; 95% CI: 68.0%, 78.8%], P = .003), whereas specificity was reduced (94.3% [829 of 879 radiographs; 95% CI: 92.8%, 95.8%], P < .001).Conclusion: A deep learning algorithm used with emergency department chest radiographs showed diagnostic performance for identifying clinically relevant abnormalities and helped improve the sensitivity of radiology residents' evaluation. Published under a CC BY 4.0 license.