Grader Variability and the Importance of Reference Standards for Evaluating Machine Learning Models for Diabetic Retinopathy

Grader Variability and the Importance of Reference Standards for Evaluating Machine Learning Models for Diabetic Retinopathy
复制标题

DOI:
10.1016/j.ophtha.2018.01.034
复制
发表时间:
2018-08-01
期刊:
影响因子:
13.7
通讯作者:
Webster, Dale R.
Webster, Dale R.
中科院分区:
医学1区
文献类型:
--
作者:
Krause, Jonathan;Gulshan, Varun;Webster, Dale R.

文献摘要

被引文献

相似文献

目的:使用裁定量化糖尿病视网膜病变(DR)分级的基础上个别graders和多数决定的错误,并训练一个改进的自动化算法DR grading.Design:回顾性analysis.Participants:视网膜眼底图像从DR筛查programmes.Methods:图像每个分级的算法,美国委员会认证的眼科医生,视网膜专家。视网膜专家的裁决共识作为参考standard.Main结果测量:不同分级者之间的协议,以及分级者和算法之间,我们测量(二次加权)Kappa评分。为了比较不同形式的手动分级的性能和各种DR严重度截止值的算法(例如,轻度或更严重的DR,中度或更严重的DR),我们测量的曲线下面积(AUC),灵敏度,和specificity.Results:193之间的裁决由视网膜专家和眼科医生的大多数决定的差异,最常见的是遗漏微动脉瘤(MA)(36%),伪影(20%),和错误分类的血管瘤(16%)。相对于参考标准,各个视网膜专家、眼科医生和算法的Kappa值范围分别为0.82 - 0.91、0.80 - 0.84和0.84。对于中度或重度DR,眼科医生的多数决定的敏感性为0.838,特异性为0.981。该算法的灵敏度为0.971,特异性为0.923,AUC为0.986。对于轻度或重度DR,该算法的灵敏度为0.970,特异性为0.917,AUC为0.986。通过使用少量裁定的共识等级作为调整数据集和更高分辨率的图像作为输入,该算法的AUC从0.934提高到0.986中度或更差的DR。结论:裁定减少了DR分级的错误。一小组裁定的DR等级允许算法性能的实质性改善。由此产生的算法的性能与美国委员会认证的眼科医生和视网膜专家的性能相当。(C)2018年美国眼科学会
Purpose: Use adjudication to quantify errors in diabetic retinopathy (DR) grading based on individual graders and majority decision, and to train an improved automated algorithm for DR grading.Design: Retrospective analysis.Participants: Retinal fundus images from DR screening programs.Methods: Images were each graded by the algorithm, U.S. board-certified ophthalmologists, and retinal specialists. The adjudicated consensus of the retinal specialists served as the reference standard.Main Outcome Measures: For agreement between different graders as well as between the graders and the algorithm, we measured the (quadratic-weighted) kappa score. To compare the performance of different forms of manual grading and the algorithm for various DR severity cutoffs (e.g., mild or worse DR, moderate or worse DR), we measured area under the curve (AUC), sensitivity, and specificity.Results: Of the 193 discrepancies between adjudication by retinal specialists and majority decision of ophthalmologists, the most common were missing microaneurysm (MAs) (36%), artifacts (20%), and misclassified hemorrhages (16%). Relative to the reference standard, the kappa for individual retinal specialists, ophthalmologists, and algorithm ranged from 0.82 to 0.91, 0.80 to 0.84, and 0.84, respectively. For moderate or worse DR, the majority decision of ophthalmologists had a sensitivity of 0.838 and specificity of 0.981. The algorithm had a sensitivity of 0.971, specificity of 0.923, and AUC of 0.986. For mild or worse DR, the algorithm had a sensitivity of 0.970, specificity of 0.917, and AUC of 0.986. By using a small number of adjudicated consensus grades as a tuning dataset and higher-resolution images as input, the algorithm improved in AUC from 0.934 to 0.986 for moderate or worse DR.Conclusions: Adjudication reduces the errors in DR grading. A small set of adjudicated DR grades allows substantial improvements in algorithm performance. The resulting algorithm's performance was on par with that of individual U.S. Board-Certified ophthalmologists and retinal specialists. (C) 2018 by the American Academy of Ophthalmology