Accuracy of Deep Learning Algorithms for the Diagnosis of Retinopathy of Prematurity by Fundus Images: A Systematic Review and Meta-Analysis.

Accuracy of Deep Learning Algorithms for the Diagnosis of Retinopathy of Prematurity by Fundus Images: A Systematic Review and Meta-Analysis.
复制标题

DOI:
10.1155/2021/8883946
复制
发表时间:
2021
影响因子:
1.9
通讯作者:
Matsuo T
Matsuo T
中科院分区:
医学4区
文献类型:
--
作者:
Zhang J;Liu Y;Mitsuhashi T;Matsuo T

文献摘要

被引文献

相似文献

早产儿视网膜病变(ROP)发生在早产儿中,可能导致失明。深度学习(DL)模型已用于眼科诊断。我们对已发表的证据进行了系统回顾和荟萃分析,以总结和评估眼底图像对ROP的DL算法诊断准确性。 我们于2021年6月13日检索了PubMed、EMBASE、Web of Science和电气与电子工程师协会Xplore数字图书馆,以寻找使用DL算法区分不同等级ROP个体的研究,这些研究提供了准确的测量结果。汇总的灵敏度和特异性值以及汇总受试者工作特征曲线(SROC)的曲线下面积(AUC)总结了总体测试性能。验证和测试数据集中的性能被一起和单独评估。对ROP的定义和等级进行了亚组分析。阈值和非阈值效应进行了测试,以评估偏差和评估与DL模型相关的准确性因素。 我们的荟萃分析包括9项研究,15个分类器。共有521,586个对象应用于DL模型。对于每项研究中的验证和测试数据集,合并的灵敏度和特异性分别为0.953(95%置信区间(CI):0.946 - 0.959)和0.975(0.973 - 0.977),AUC为0.984(0.978 - 0.989)。对于验证数据集和测试数据集,AUC分别为0.977(0.968 - 0.986)和0.987(0.982 - 0.992)。在ROP与正常和两个ROP等级分化的亚组分析中,AUC分别为0.990(0.944 - 0.994)和0.982(0.964 - 0.999)。 我们的研究表明,DL模型可以发挥重要作用,在检测和分级ROP具有高灵敏度,特异性和可重复性。基于DL的自动化系统的应用可能会在未来改善ROP的筛查和诊断。
Retinopathy of prematurity (ROP) occurs in preterm infants and may contribute to blindness. Deep learning (DL) models have been used for ophthalmologic diagnoses. We performed a systematic review and meta-analysis of published evidence to summarize and evaluate the diagnostic accuracy of DL algorithms for ROP by fundus images. We searched PubMed, EMBASE, Web of Science, and Institute of Electrical and Electronics Engineers Xplore Digital Library on June 13, 2021, for studies using a DL algorithm to distinguish individuals with ROP of different grades, which provided accuracy measurements. The pooled sensitivity and specificity values and the area under the curve (AUC) of summary receiver operating characteristics curves (SROC) summarized overall test performance. The performances in validation and test datasets were assessed together and separately. Subgroup analyses were conducted between the definition and grades of ROP. Threshold and nonthreshold effects were tested to assess biases and evaluate accuracy factors associated with DL models. Nine studies with fifteen classifiers were included in our meta-analysis. A total of 521,586 objects were applied to DL models. For combined validation and test datasets in each study, the pooled sensitivity and specificity were 0.953 (95% confidence intervals (CI): 0.946–0.959) and 0.975 (0.973–0.977), respectively, and the AUC was 0.984 (0.978–0.989). For the validation dataset and test dataset, the AUC was 0.977 (0.968–0.986) and 0.987 (0.982–0.992), respectively. In the subgroup analysis of ROP vs. normal and differentiation of two ROP grades, the AUC was 0.990 (0.944–0.994) and 0.982 (0.964–0.999), respectively. Our study shows that DL models can play an essential role in detecting and grading ROP with high sensitivity, specificity, and repeatability. The application of a DL-based automated system may improve ROP screening and diagnosis in the future.