Low-Shot Deep Learning of Diabetic Retinopathy With Potential Applications to Address Artificial Intelligence Bias in Retinal Diagnostics and Rare Ophthalmic Diseases

Low-Shot Deep Learning of Diabetic Retinopathy With Potential Applications to Address Artificial Intelligence Bias in Retinal Diagnostics and Rare Ophthalmic Diseases
复制标题

DOI:
10.1001/jamaophthalmol.2020.3269
复制
发表时间:
2020-10-01
期刊:
影响因子:
8.1
通讯作者:
Bressler, Neil M.
Bressler, Neil M.
中科院分区:
医学1区
文献类型:
--
作者:
Burlina, Philippe;Paul, William;Bressler, Neil M.

文献摘要

被引文献

相似文献

最近的研究已经证明了人工智能(AI)在自动视网膜疾病诊断中的成功应用,但还没有解决深度学习系统面临的一个根本挑战:目前需要大量标准注释的视网膜数据集进行训练。旨在从相对少量的训练数据中学习的低激发学习算法可能有益于涉及罕见视网膜疾病的临床情况,或者当解决由可能不足以代表某些训练组的数据引起的潜在偏差时,例如85岁以上的个体。目的评估低-当使用小的训练数据集进行自动视网膜诊断时,射击深度学习方法是有益的。设计、设置和参与者这项横断面研究于2019年7月1日至6月21日进行,2020年,比较了不同的糖尿病视网膜病变分类算法,传统和低射,用于2类指定(糖尿病视网膜病变转诊与未转诊)。使用公共领域EyePACS数据集,其中最初包括来自44346人的88692个基金。统计分析时间为2020年2月1日至6月21日。主要结果和指标通过受试者工作曲线及其曲线下面积(AUC)、精确召回曲线、准确率和F1评分来衡量各种AI算法的性能(95%CI),并针对不同的训练数据大小进行评估,范围从5120到10个样本每类。结果深度学习算法,当训练足够大的数据集时(每类5120个样本),产生了相当的性能,AUC为0.8330(95% CI,0.8140-0.8520)传统方法(例如,微调ResNet),与低激发方法(AUC,0.8348 [95% CI,0.8159-0.8537])(使用自我监督的Deep InfoMax [我们的方法表示为DIM])相比。然而,当可用的训练图像少得多时(n = 160),传统的深度学习方法的AUC降至0.6585(95%CI,0.6332-0.6838),并且通过使用自我监督的低拍摄方法的AUC为0.7467(95%CI,0.7239-0.7695)。在非常低的射击(n = 10),传统方法的性能接近机会,AUC为0.5178(95% CI,0.4909-0.5447)与最佳低激发方法相比(AUC,0.5778 [95%CI,0.5512-0.6044])。结论和相关性这些发现表明,使用低浓度的当有限数量的注释训练视网膜图像可用时(例如,罕见的眼科疾病或解决潜在的AI偏见时),AI视网膜诊断的拍摄方法。
IMPORTANCE Recent studies have demonstrated the successful application of artificial intelligence (AI) for automated retinal disease diagnostics but have not addressed a fundamental challenge for deep learning systems: the current need for large, criterion standard-annotated retinal data sets for training. Low-shot learning algorithms, aiming to learn from a relatively low number of training data, may be beneficial for clinical situations involving rare retinal diseases or when addressing potential bias resulting from data that may not adequately represent certain groups for training, such as individuals older than 85 years.OBJECTIVE To evaluate whether low-shot deep learning methods are beneficial when using small training data sets for automated retinal diagnostics.DESIGN, SETTING, AND PARTICIPANTS This cross-sectional study, conducted from July 1, 2019, to June 21, 2020, compared different diabetic retinopathy classification algorithms, traditional and low-shot, for 2-class designations (diabetic retinopathy warranting referral vs not warranting referral). The public domain EyePACS data set was used, which originally included 88 692 fundi from 44 346 individuals. Statistical analysis was performed from February 1 to June 21, 2020.MAIN OUTCOMES AND MEASURES The performance (95% CIs) of the various AI algorithms was measured via receiver operating curves and their area under the curve (AUC), precision recall curves, accuracy, and F1 score, evaluated for different training data sizes, ranging from 5120 to 10 samples per class.RESULTS Deep learning algorithms, when trained with sufficiently large data sets (5120 samples per class), yielded comparable performance, with an AUC of 0.8330 (95% CI, 0.8140-0.8520) for a traditional approach (eg, fined-tuned ResNet), compared with low-shot methods (AUC, 0.8348 [95% CI, 0.8159-0.8537]) (using self-supervised Deep InfoMax [our method denoted as DIM]). However, when far fewer training images were available (n = 160), the traditional deep learning approach had an AUC decreasing to 0.6585 (95% CI, 0.6332-0.6838) and was outperformed by a low-shot method using self-supervision with an AUC of 0.7467 (95% CI, 0.7239-0.7695). At very low shots (n = 10), the traditional approach had performance close to chance, with an AUC of 0.5178 (95% CI, 0.4909-0.5447) compared with the best low-shot method (AUC, 0.5778 [95% CI, 0.5512-0.6044]).CONCLUSIONS AND RELEVANCE These findings suggest the potential benefits of using low-shot methods for AI retinal diagnostics when a limited number of annotated training retinal images are available (eg, with rare ophthalmic diseases or when addressing potential AI bias).