Effect of a comprehensive deep-learning model on the accuracy of chest x-ray interpretation by radiologists: a retrospective, multireader multicase study

Effect of a comprehensive deep-learning model on the accuracy of chest x-ray interpretation by radiologists: a retrospective, multireader multicase study
复制标题

DOI:
10.1016/s2589-7500(21)00106-0
复制
发表时间:
2021-07-26
影响因子:
30.8
通讯作者:
Jones, Catherine M.
Jones, Catherine M.
中科院分区:
医学1区
文献类型:
--
作者:
Seah, Jarrel C. Y.;Tang, Cyril H. M.;Jones, Catherine M.

文献摘要

被引文献

相似文献

背景胸部X线检查广泛应用于临床实践;然而,由于人为错误和缺乏经验丰富的胸部放射科医生,解释可能会受到阻碍。深度学习有可能提高胸部X射线解释的准确性。因此,我们的目标是评估放射科医生在有和没有深度学习模型的帮助下的准确性。在这项回顾性研究中,我们在来自澳大利亚、欧洲和美国的5个数据集的821 681张图像(284 649名患者)上训练了一个深度学习模型。测试数据集中包括2568例至少接受过一次正面胸部X线检查的成年患者(≥ 16岁)的丰富胸部X线检查病例;病例代表住院、门诊和急诊情况。20名放射科医生在3个月的洗脱期内审查了有和没有深度学习模型帮助的病例。我们评估了使用深度学习模型作为决策支持时,127个临床结果的胸部X射线解释准确性的变化,方法是计算每个放射科医生在使用和不使用深度学习模型时的受试者工作特征曲线下面积(AUC)。我们还比较了单独模型的AUC与无辅助放射科医生的AUC。如果模型和无辅助放射科医生之间AUC差异的校正后95% CI下限大于-0middot05,则认为模型在该结果方面具有非劣效性。如果下限超过0,则认为该模型具有上级性。在127个临床结果中,无辅助放射科医生的宏观平均AUC为0 middot 713(95%CI 0 middot 645 - 0 middot 785),而在模型辅助下为0 middot 808(0 middot 763 - 0 middot 839)。深度学习模型在统计学上显著提高了放射科医生对127个临床结果中的102个(80%)的分类准确性,在统计学上不劣于19个(15%)结果,并且当放射科医生使用深度学习模型时,没有发现准确性下降。无辅助放射科医生在所有发现中的宏观平均AUC为0 middot 713(0 middot 645 - 0 middot 785),而单独模型的AUC为0 middot 957(0 middot 954 - 0 middot 959)。在模型预测的124项临床结果中,单独的模型分类比无辅助放射科医生的117项(94%)更准确,并且在所有其他临床结果中不劣于无辅助放射科医生。解释本研究显示了综合深度学习模型在广泛的临床实践中改善胸部X射线解释的潜力。资助Annalise.ai.版权所有(c)2021作者。由Elsevier Ltd.出版。这是一篇开放获取文章,使用CC BY 4.0许可证。
Background Chest x-rays are widely used in clinical practice; however, interpretation can be hindered by human error and a lack of experienced thoracic radiologists. Deep learning has the potential to improve the accuracy of chest x-ray interpretation. We therefore aimed to assess the accuracy of radiologists with and without the assistance of a deep learning model. Methods In this retrospective study, a deep-learning model was trained on 821 681 images (284 649 patients) from five data sets from Australia, Europe, and the USA. 2568 enriched chest x-ray cases from adult patients (>= 16 years) who had at least one frontal chest x-ray were included in the test dataset; cases were representative of inpatient, outpatient, and emergency settings. 20 radiologists reviewed cases with and without the assistance of the deep-learning model with a 3-month washout period. We assessed the change in accuracy of chest x-ray interpretation across 127 clinical findings when the deep-learning model was used as a decision support by calculating area under the receiver operating characteristic curve (AUC) for each radiologist with and without the deep-learning model. We also compared AUCs for the model alone with those of unassisted radiologists. If the lower bound of the adjusted 95% CI of the difference in AUC between the model and the unassisted radiologists was more than -0middot05, the model was considered to be non-inferior for that finding. If the lower bound exceeded 0, the model was considered to be superior. Findings Unassisted radiologists had a macroaveraged AUC of 0middot713 (95% CI 0middot645-0middot785) across the 127 clinical findings, compared with 0middot808 (0middot763-0middot839) when assisted by the model. The deep-learning model statistically significantly improved the classification accuracy of radiologists for 102 (80%) of 127 clinical findings, was statistically non-inferior for 19 (15%) findings, and no findings showed a decrease in accuracy when radiologists used the deep learning model. Unassisted radiologists had a macroaveraged mean AUC of 0middot713 (0middot645-0middot785) across all findings, compared with 0middot957 (0middot954-0middot959) for the model alone. Model classification alone was significantly more accurate than unassisted radiologists for 117 (94%) of 124 clinical findings predicted by the model and was non-inferior to unassisted radiologists for all other clinical findings. Interpretation This study shows the potential of a comprehensive deep-learning model to improve chest x-ray interpretation across a large breadth of clinical practice. Funding Annalise.ai. Copyright (c) 2021 The Author(s). Published by Elsevier Ltd. This is an Open Access article under the CC BY 4.0 license.