Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis.

Diagnostic accuracy of deep learning in medical imaging: a systematic review and meta-analysis.
复制标题

DOI:
10.1038/s41746-021-00438-z
复制
发表时间:
2021-04-07
影响因子:
15.2
通讯作者:
Darzi A
Darzi A
中科院分区:
医学1区
文献类型:
--
作者:
Aggarwal R;Sounderajah V;Martin G;Ting DSW;Karthikesalingam A;King D;Ashrafian H;Darzi A

文献摘要

参考文献

被引文献

相似文献

深度学习具有改变医学诊断学的潜力。然而,DL的诊断准确性尚不确定。我们的目的是评估DL算法在医学成像中识别病理的诊断准确性。截至2020年1月,在Medline和EMBASE上进行了搜索。我们确定了11,921项研究,其中503项纳入了系统评价。纳入荟萃分析的眼科研究有82篇,乳房疾病有82篇,呼吸系统疾病有115篇。其他专业的224项研究纳入定性综述。报告了使用医学成像识别病理的DL算法的诊断准确性的同行评议研究被包括在内。主要结果是文献中对诊断准确性、研究设计和报告标准的测量。使用随机效应荟萃分析合并估计。在眼科,视网膜眼底照片和光学相干断层扫描诊断糖尿病视网膜病变、老年性黄斑变性和青光眼的AUC在0.933到1之间。在呼吸成像方面,胸部X线或CT扫描诊断肺结节或肺癌的AUC值在0.864至0.937之间。对于乳房成像,乳房X光检查、超声、核磁共振和数字乳腺断层扫描诊断乳腺癌的AUC值在0.868到0.909之间。研究之间的异质性很高,注意到在方法、术语和结果测量方面存在着广泛的差异。这可能会导致对医学成像上的DL算法诊断精度的高估。迫切需要制定专门针对人工智能的赤道准则,特别是STARD准则,以便就该领域的关键问题提供指导。
Deep learning (DL) has the potential to transform medical diagnostics. However, the diagnostic accuracy of DL is uncertain. Our aim was to evaluate the diagnostic accuracy of DL algorithms to identify pathology in medical imaging. Searches were conducted in Medline and EMBASE up to January 2020. We identified 11,921 studies, of which 503 were included in the systematic review. Eighty-two studies in ophthalmology, 82 in breast disease and 115 in respiratory disease were included for meta-analysis. Two hundred twenty-four studies in other specialities were included for qualitative review. Peer-reviewed studies that reported on the diagnostic accuracy of DL algorithms to identify pathology using medical imaging were included. Primary outcomes were measures of diagnostic accuracy, study design and reporting standards in the literature. Estimates were pooled using random-effects meta-analysis. In ophthalmology, AUC’s ranged between 0.933 and 1 for diagnosing diabetic retinopathy, age-related macular degeneration and glaucoma on retinal fundus photographs and optical coherence tomography. In respiratory imaging, AUC’s ranged between 0.864 and 0.937 for diagnosing lung nodules or lung cancer on chest X-ray or CT scan. For breast imaging, AUC’s ranged between 0.868 and 0.909 for diagnosing breast cancer on mammogram, ultrasound, MRI and digital breast tomosynthesis. Heterogeneity was high between studies and extensive variation in methodology, terminology and outcome measures was noted. This can lead to an overestimation of the diagnostic accuracy of DL algorithms on medical imaging. There is an immediate need for the development of artificial intelligence-specific EQUATOR guidelines, particularly STARD, in order to provide guidance around key issues in this field.
DOI: 10.1016/j.ogla.2019.03.008
发表时间: 2019-07-01
影响因子: 2.9
作者:
Asaoka, Ryo;Tanito, Masaki;Kiuchi, Yoshiaki
通讯作者: Kiuchi, Yoshiaki
Stard 2015:用于报告诊断准确性研究的基本项目的更新列表。
DOI: 10.1136/bmj.h5527
发表时间: 2015-10-28
期刊: BMJ (Clinical research ed.)
影响因子: --
作者:
Bossuyt PM;Reitsma JB;Bruns DE;Gatsonis CA;Glasziou PP;Irwig L;Lijmer JG;Moher D;Rennie D;de Vet HC;Kressel HY;Rifai N;Golub RM;Altman DG;Hooft L;Korevaar DA;Cohen JF;STARD Group
通讯作者: STARD Group
DOI: 10.1097/ijg.0000000000001319
发表时间: 2019-12-01
影响因子: 2
作者:
Al-Aswad, Lama A.;Kapoor, Rahul;Moazami, Golnaz
通讯作者: Moazami, Golnaz
DOI: 10.1016/s2589-7500(19)30004-4
发表时间: 2019-05-01
影响因子: 30.8
作者:
Bellemo, Valentina;Lim, Zhan W.;Ting, Daniel S. W.
通讯作者: Ting, Daniel S. W.
DOI: 10.1007/s11517-019-02066-y
发表时间: 2019-11-28
影响因子: 3.2
作者:
Alqudah, Ali Mohammad
通讯作者: Alqudah, Ali Mohammad