Performance of a Deep Learning Model vs Human Reviewers in Grading Endoscopic Disease Severity of Patients With Ulcerative Colitis

Performance of a Deep Learning Model vs Human Reviewers in Grading Endoscopic Disease Severity of Patients With Ulcerative Colitis
复制标题

DOI:
10.1001/jamanetworkopen.2019.3963
复制
发表时间:
2019-05-01
期刊:
影响因子:
13.8
通讯作者:
Waljee, Akbar K.
Waljee, Akbar K.
中科院分区:
医学1区
文献类型:
--
作者:
Stidham, Ryan W.;Liu, Wenshuo;Waljee, Akbar K.

文献摘要

被引文献

相似文献

重要性评估溃疡性结肠炎(UC)的内镜疾病严重程度是确定治疗反应的关键因素,但其在临床实践中的使用受到经验丰富的人类评审员要求的限制。目的确定深度学习模型是否可以对UC的内镜严重程度以及经验丰富的人类评审员进行分级。设计、设置和参与者在这项诊断研究中,使用4级马约子评分对内窥镜图像进行回顾性分级,由2名独立的评审员进行,评分差异由第3名评审员裁定。使用2007年1月1日至2017年12月31日期间在美国一家三级护理转诊中心接受结肠镜检查的3082名UC患者的16514张图像,构建了一个159层卷积神经网络(CNN)作为深度学习模型,以训练图像并将其分类为2个临床相关组:缓解(马约分项评分0或1)和中度至重度疾病(马约分项评分,2或3)。90%的队列用于构建模型,10%用于测试模型;该过程重复10次。一组30个全运动结肠镜视频,看不见的模型,然后用于外部验证,以模仿现实世界的application.Main结果和测量模型的性能进行了评估,使用的受试者工作曲线下面积(AUROC),灵敏度和特异性,阳性预测值(PPV),和阴性预测值(NPV)。Kappa统计量(kappa)用于衡量CNN相对于裁定的人类参考核心的一致性。结果作者纳入了来自3082名独特患者的16514张图像(中位[IQR]年龄,41.3 [26.1-61.8]岁,1678 [54.4%]女性),其中3980张图像(24.1%)被判定的参考评分归类为中度至重度疾病。CNN在区分内镜缓解与中度至重度疾病方面表现出色,AUROC为0.966(95% CI,0.967-0.972); PPV为0.87(95% CI,0.85-0.88),灵敏度为83.0%(95%CI,80.8%-85.4%),特异性为96.0%(95%CI,95.1%-97.1%); NPV为0.94(95%CI,0.93-0.95)。CNN与裁定的参考评分之间的加权kappa一致性也有利于识别准确的马约子评分(kappa = 0.84; 95% CI,0.83-0.86),并且与经验丰富的评审员之间的一致性相似(kappa = 0.86; 95% CI,0.85-0.87)。将CNN应用于整个结肠镜检查视频对于识别中度至重度疾病具有相似的准确性(AUROC,0.97; 95%CI,0.963-0.969)。结论和相关性本研究发现,深度学习模型的性能与有经验的人类评审员在对UC的内窥镜严重程度进行分级方面相似。鉴于其可扩展性,这种方法可以提高结肠镜检查在研究和常规实践中对UC的使用。
IMPORTANCE Assessing endoscopic disease severity in ulcerative colitis (UC) is a key element in determining therapeutic response, but its use in clinical practice is limited by the requirement for experienced human reviewers.OBJECTIVE To determine whether deep learning models can grade the endoscopic severity of UC as well as experienced human reviewers.DESIGN, SETTING, AND PARTICIPANTS In this diagnostic study, retrospective grading of endoscopic images using the 4-level Mayo subscore was performed by 2 independent reviewers with score discrepancies adjudicated by a third reviewer. Using 16 514 images from 3082 patients with UC who underwent colonoscopy at a single tertiary care referral center in the United States between January 1, 2007, and December 31, 2017, a 159-layer convolutional neural network (CNN) was constructed as a deep learning model to train and categorize images into 2 clinically relevant groups: remission (Mayo subscore 0 or 1) and moderate to severe disease (Mayo subscore, 2 or 3). Ninety percent of the cohort was used to build the model and 10% was used to test it; the process was repeated 10 times. A set of 30 full-motion colonoscopy videos, unseen by the model, was then used for external validation to mimic real-world application.MAIN OUTCOMES AND MEASURES Model performance was assessed using area under the receiver operating curve (AUROC), sensitivity and specificity, positive predictive value (PPV), and negative predictive value (NPV). Kappa statistics (kappa) were used to measure agreement of the CNN relative to adjudicated human reference cores.RESULTS The authors included 16 514 images from 3082 unique patients (median [IQR] age, 41.3 [26.1-61.8] years, 1678 [54.4%] female), with 3980 images (24.1%) classified as moderate-to-severe disease by the adjudicated reference score. The CNN was excellent for distinguishing endoscopic remission from moderate-to-severe disease with an AUROC of 0.966 (95% CI, 0.967-0.972); a PPV of 0.87 (95% CI, 0.85-0.88) with a sensitivity of 83.0% (95% CI, 80.8%-85.4%) and specificty of 96.0% (95% CI, 95.1%-97.1%); and NPV of 0.94 (95% CI, 0.93-0.95). Weighted kappa agreement between the CNN and the adjudicated reference score was also good for identifying exact Mayo subscores (kappa = 0.84; 95% CI, 0.83-0.86) and was similar to the agreement between experienced reviewers (kappa = 0.86; 95% CI, 0.85-0.87). Applying the CNN to entire colonoscopy videos had similar accuracy for identifying moderate to severe disease (AUROC, 0.97; 95% CI, 0.963-0.969).CONCLUSIONS AND RELEVANCE This study found that deep learning model performance was similar to experienced human reviewers in grading endoscopic severity of UC. Given its scalability, this approach could improve the use of colonoscopy for UC in both research and routine practice.