Evaluation of a Deep Learning System For Identifying Glaucomatous Optic Neuropathy Based on Color Fundus Photographs

Evaluation of a Deep Learning System For Identifying Glaucomatous Optic Neuropathy Based on Color Fundus Photographs
复制标题

DOI:
10.1097/ijg.0000000000001319
复制
发表时间:
2019-12-01
影响因子:
2
通讯作者:
Moazami, Golnaz
Moazami, Golnaz
中科院分区:
医学3区
文献类型:
--
作者:
Al-Aswad, Lama A.;Kapoor, Rahul;Moazami, Golnaz

文献摘要

被引文献

相似文献

简述:Pegasus 在诊断表现方面优于 6 名眼科医生中的 5 名,并且深度学习系统与眼科医生之间的“最佳案例”共识之间没有统计学上的显着差异。 Pegasus 与金标准的一致性为 0.715,而眼科医生与金标准的最高一致性为 0.613。此外,Pegasus 的高灵敏度使其成为筛查青光眼视神经病变患者的宝贵工具。目的:本研究的目的是评估深度学习系统在识别青光眼视神经病变方面的性能。材料和方法:在这项回顾性单中心研究中,六位眼科医生和深度学习系统 Pegasus 对 110 张彩色眼底照片进行了分级。患者图像是从新加坡马来眼科研究中随机抽取的。眼科医生和 Pegasus 进行了相互比较,并与新加坡马来眼科研究给出的原始临床诊断进行了比较,该诊断被定义为金标准。 Pegasus 的表现与“最佳情况”共识场景进行了比较,该场景是眼科医生的共识意见与黄金标准最接近的组合。眼科医生和 Pegasus 在根据眼底照片对非青光眼与青光眼进行二元分类时,根据敏感性、特异性和受试者工作特征曲线下面积 (AUROC) 进行评估,并确定观察者内和观察者间的一致性。结果:Pegasus 实现了 92.6% 的 AUROC,而眼科医生的 AUROC 范围为 69.6% 至 84.9%,“最佳情况”共识情景 AUROC 为 89.1%。 Pegasus 的敏感性为 83.7%,特异性为 88.2%,而眼科医生的敏感性范围为 61.3% 至 81.6%,特异性范围为 80.0% 至 94.1%。 Pegasus 与金标准的一致性为 0.715,而眼科医生与金标准的最高一致性为 0.613。眼科医生的观察者内一致性范围为 0.62 到 0.97,Pegasus 的一致性非常好(1.00)。深度学习系统花费了眼科医生 10% 的时间来确定分类。结论:Pegasus 在诊断性能方面优于 6 名眼科医生中的 5 名,并且深度学习系统与眼科医生之间的“最佳案例”共识之间没有统计学上的显着差异。 Pegasus 的高灵敏度使其成为筛查青光眼视神经病变患者的宝贵工具。未来的工作将把这项研究扩展到更大的患者样本。
Precis: Pegasus outperformed 5 of the 6 ophthalmologists in terms of diagnostic performance, and there was no statistically significant difference between the deep learning system and the "best case" consensus between the ophthalmologists. The agreement between Pegasus and gold standard was 0.715, whereas the highest ophthalmologist agreement with the gold standard was 0.613. Furthermore, the high sensitivity of Pegasus makes it a valuable tool for screening patients with glaucomatous optic neuropathy. Purpose: The purpose of this study was to evaluate the performance of a deep learning system for the identification of glaucomatous optic neuropathy. Materials and Methods: Six ophthalmologists and the deep learning system, Pegasus, graded 110 color fundus photographs in this retrospective single-center study. Patient images were randomly sampled from the Singapore Malay Eye Study. Ophthalmologists and Pegasus were compared with each other and to the original clinical diagnosis given by the Singapore Malay Eye Study, which was defined as the gold standard. Pegasus' performance was compared with the "best case" consensus scenario, which was the combination of ophthalmologists whose consensus opinion most closely matched the gold standard. The performance of the ophthalmologists and Pegasus, at the binary classification of nonglaucoma versus glaucoma from fundus photographs, was assessed in terms of sensitivity, specificity and the area under the receiver operating characteristic curve (AUROC), and the intraobserver and interobserver agreements were determined. Results: Pegasus achieved an AUROC of 92.6% compared with ophthalmologist AUROCs that ranged from 69.6% to 84.9% and the "best case" consensus scenario AUROC of 89.1%. Pegasus had a sensitivity of 83.7% and a specificity of 88.2%, whereas the ophthalmologists' sensitivity ranged from 61.3% to 81.6% and specificity ranged from 80.0% to 94.1%. The agreement between Pegasus and gold standard was 0.715, whereas the highest ophthalmologist agreement with the gold standard was 0.613. Intraobserver agreement ranged from 0.62 to 0.97 for ophthalmologists and was perfect (1.00) for Pegasus. The deep learning system took similar to 10% of the time of the ophthalmologists in determining classification. Conclusions: Pegasus outperformed 5 of the 6 ophthalmologists in terms of diagnostic performance, and there was no statistically significant difference between the deep learning system and the "best case" consensus between the ophthalmologists. The high sensitivity of Pegasus makes it a valuable tool for screening patients with glaucomatous optic neuropathy. Future work will extend this study to a larger sample of patients.