Diagnosis of thyroid cancer using deep convolutional neural network models applied to sonographic images: a retrospective, multicohort, diagnostic study

Diagnosis of thyroid cancer using deep convolutional neural network models applied to sonographic images: a retrospective, multicohort, diagnostic study
复制标题

使用应用于超声图像的深度卷积神经网络模型诊断甲状腺癌:一项回顾性、多队列诊断研究

DOI:
10.1016/s1470-2045(18)30762-9
复制
发表时间:
2019-02-01
期刊:
影响因子:
51.1
通讯作者:
Chen, Kexin
Chen, Kexin
中科院分区:
医学1区
文献类型:
--
作者:
Li, Xiangchun;Zhang, Sheng;Chen, Kexin

文献摘要

被引文献

相似文献

背景:由于广泛使用敏感成像技术进行筛查,导致过度诊断和过度治疗,甲状腺癌的发病率正在稳步上升。这种总体发病率的增长主要是由于惰性和分化良好的乳头状亚型和早期甲状腺癌的诊断增加,而晚期甲状腺癌的发病率略有增加。甲状腺超声常用于诊断甲状腺癌。本研究的目的是利用深度卷积神经网络(DCNN)模型通过分析临床超声图像数据来提高甲状腺癌的诊断准确性。方法利用中国三家医院的超声图像集进行回顾性、多队列、诊断性研究。我们在天津市肿瘤医院甲状腺影像数据库中的17 627例甲状腺癌患者的131 731张超声图像和25 325例对照的180 668张图像的训练集上开发并训练了DCNN模型。训练集的临床诊断由天津市肿瘤医院的16名放射科医师进行。从解剖部位被判断为没有癌症的图像被排除在训练集之外,只有疑似甲状腺癌的个体进行病理检查以确认诊断。模型的诊断性能在来自天津肿瘤医院的内部验证集(来自1118名患者的8606张图像)和中国的两个外部数据集(吉林中西医结合医院,来自154名患者的741张图像;山东威海市立医院,来自1420名患者的11039张图像)中进行了验证。验证组中所有临床检查后疑似甲状腺癌的个体均行病理检查。我们还将DCNN模型的特异性和敏感性与六位熟练甲状腺超声放射科医生在三个验证集上的表现进行了比较。在2012年1月1日至2018年3月28日期间,获得了四个研究队列的超声图像。模型对甲状腺癌患者的识别效果较好,天津内部验证集的曲线下面积为0.947 (95% CI 0.935-0.959),吉林外部验证集的曲线下面积为0.912 (95% CI 0.865-0.958),威海外部验证集的曲线下面积为0.908 (95% CI 0.891-0.925)。与熟练的放射科医生相比,DCNN模型在识别甲状腺癌患者方面也表现出更好的表现。对于天津内部验证集,敏感性为93.4% (95% CI 89.6-96.1)对96.9% (93.9-98.6,p=0.003),特异性为86.1%(81.1-90.2)对59.4% (53.0-65.6,p< 0.0001)。对于吉林外部验证集,敏感性为84.3% (95% CI 73.6-91.9)对92.9% (84.1-97.6,p=0.048),特异性为86.9% (95% CI 77.8-93.3)对57.1% (45.9-67.9,p< 0.0001)。对于威海外部验证集,敏感性为84.7% (95% CI 77.0 ~ 90.7) vs 89.0% (81.9 ~ 94.0, p=0.25),特异性为87.8% (95% CI 81.6 ~ 92.5) vs 68.6% (60.7 ~ 75.8, p< 0.0001)。与一组熟练的放射科医生相比,DCNN模型在识别甲状腺癌患者方面表现出相似的敏感性和更高的特异性。作为随机临床试验的一部分,DCNN模型的改进技术性能值得进一步研究。爱思唯尔有限公司版权所有版权所有。
Background The incidence of thyroid cancer is rising steadily because of overdiagnosis and overtreatment conferred by widespread use of sensitive imaging techniques for screening. This overall incidence growth is especially driven by increased diagnosis of indolent and well-differentiated papillary subtype and early-stage thyroid cancer, whereas the incidence of advanced-stage thyroid cancer has increased marginally. Thyroid ultrasound is frequently used to diagnose thyroid cancer. The aim of this study was to use deep convolutional neural network (DCNN) models to improve the diagnostic accuracy of thyroid cancer by analysing sonographic imaging data from clinical ultrasounds.Methods We did a retrospective, multicohort, diagnostic study using ultrasound images sets from three hospitals in China. We developed and trained the DCNN model on the training set, 131 731 ultrasound images from 17 627 patients with thyroid cancer and 180 668 images from 25 325 controls from the thyroid imaging database at Tianjin Cancer Hospital. Clinical diagnosis of the training set was made by 16 radiologists from Tianjin Cancer Hospital. Images from anatomical sites that were judged as not having cancer were excluded from the training set and only individuals with suspected thyroid cancer underwent pathological examination to confirm diagnosis. The model's diagnostic performance was validated in an internal validation set from Tianjin Cancer Hospital (8606 images from 1118 patients) and two external datasets in China (the Integrated Traditional Chinese and Western Medicine Hospital, Jilin, 741 images from 154 patients; and the Weihai Municipal Hospital, Shandong, 11 039 images from 1420 patients). All individuals with suspected thyroid cancer after clinical examination in the validation sets had pathological examination. We also compared the specificity and sensitivity of the DCNN model with the performance of six skilled thyroid ultrasound radiologists on the three validation sets.Findings Between Jan 1, 2012, and March 28, 2018, ultrasound images for the four study cohorts were obtained. The model achieved high performance in identifying thyroid cancer patients in the validation sets tested, with area under the curve values of 0.947 (95% CI 0.935-0.959) for the Tianjin internal validation set, 0.912 (95% CI 0.865-0.958) for the Jilin external validation set, and 0.908 (95% CI 0.891-0.925) for the Weihai external validation set. The DCNN model also showed improved performance in identifying thyroid cancer patients versus skilled radiologists. For the Tianjin internal validation set, sensitivity was 93.4% (95% CI 89.6-96.1) versus 96.9% (93.9-98.6; p=0.003) and specificity was 86.1% (81.1-90.2) versus 59.4% (53.0-65.6; p< 0.0001). For the Jilin external validation set, sensitivity was 84.3% (95% CI 73.6-91.9) versus 92.9% (84.1-97.6; p=0.048) and specificity was 86.9% (95% CI 77.8-93.3) versus 57.1% (45.9-67.9; p< 0.0001). For the Weihai external validation set, sensitivity was 84.7% (95% CI 77.0-90.7) versus 89.0% (81.9-94.0; p=0.25) and specificity was 87.8% (95% CI 81.6-92.5) versus 68.6% (60.7-75.8; p< 0.0001).Interpretation The DCNN model showed similar sensitivity and improved specificity in identifying patients with thyroid cancer compared with a group of skilled radiologists. The improved technical performance of the DCNN model warrants further investigation as part of randomised clinical trials. Copyright (C) 2018 Elsevier Ltd. All rights reserved.