Classification of Cancer at Prostate MRI: Deep Learning versus Clinical PI-RADS Assessment

Classification of Cancer at Prostate MRI: Deep Learning versus Clinical PI-RADS Assessment
复制标题

DOI:
10.1148/radiol.2019190938
复制
发表时间:
2019-12-01
期刊:
影响因子:
19.7
通讯作者:
Bonekamp, David
Bonekamp, David
中科院分区:
医学1区
文献类型:
--
作者:
Schelb, Patrick;Kohl, Simon;Bonekamp, David

文献摘要

被引文献

相似文献

背景:越来越多怀疑患有临床显著前列腺癌(sPC)的男性接受前列腺MRI检查。深度学习为人类解释提供诊断支持的潜力需要进一步评估。目的:比较临床评估与基于t2加权和弥散MRI训练的深度学习分割系统在检测和分割sPC可疑病变方面的性能。材料与方法:在本回顾性研究中,对2015年至2016年使用单一3.0 t MRI系统检查的连续男性的t2加权和弥散性前列腺MRI序列进行人工分割。通过联合靶向和扩展系统mri经直肠US融合活检提供了基本事实,sPC定义为国际泌尿病理学学会Gleason分级组大于或等于2。通过使用分样本验证,U-Net通过交叉验证在训练集(80%的数据)上进行内部验证,随后在测试集(20%的数据)上进行外部验证。通过将基于六分仪的交叉验证性能与前列腺成像报告和数据系统(PI-RADS)的临床性能相匹配,对u - net衍生的sPC概率图进行校准。通过灵敏度、特异性、预测值和Dice系数对PI-RADS和U-Net的性能进行比较。结果:共评估312名男性(中位年龄64岁,四分位间距[IQR]为58-71岁)。训练集包括250名男性(中位年龄64岁,IQR为58-71岁),测试集包括62名男性(中位年龄64岁,IQR为60-69岁)。在测试集中,PI-RADS在每个患者基础上大于或等于3的截止点与大于或等于4的截止点的敏感性分别为96%(26人中25人)和88%(26人中23人),特异性分别为22%(36人中8人)和50%(36人中18人)。U-Net在概率阈值大于等于0.22和大于等于0.33时的敏感性分别为96%(25 / 26)和92%(24 / 26)(均为P>.99),特异性分别为31%(11 / 36)和47%(17 / 36)(均为P>.99),与PI-RADS无统计学差异。前列腺的Dice系数为0.89,MRI病变分割的Dice系数为0.35。在测试集中,PI-RADS大于或等于4与U-Net病变的符合性将U-Net概率阈值大于或等于0.33的阳性预测值从48%(58人中的28人)提高到67%(36人中的24人)(P = 0.01),而阴性预测值保持不变(83%[30人中的25人]vs 83%[52人中的43人];P = 0.99)。结论:经t2加权和弥散MRI训练的U-Net达到了与临床前列腺影像学报告和数据系统评估相似的性能。(c) rsna, 2019
Background: Men suspected of having clinically significant prostate cancer (sPC) increasingly undergo prostate MRI. The potential of deep learning to provide diagnostic support for human interpretation requires further evaluation.Purpose: To compare the performance of clinical assessment to a deep learning system optimized for segmentation trained with T2-weighted and diffusion MRI in the task of detection and segmentation of lesions suspicious for sPC.Materials and Methods: In this retrospective study, T2-weighted and diffusion prostate MRI sequences from consecutive men examined with a single 3.0-T MRI system between 2015 and 2016 were manually segmented. Ground truth was provided by combined targeted and extended systematic MRI-transrectal US fusion biopsy, with sPC defined as International Society of Urological Pathology Gleason grade group greater than or equal to 2. By using split-sample validation, U-Net was internally validated on the training set (80% of the data) through cross validation and subsequently externally validated on the test set (20% of the data). U-Net-derived sPC probability maps were calibrated by matching sextant-based cross-validation performance to clinical performance of Prostate Imaging Reporting and Data System (PI-RADS). Performance of PI-RADS and U-Net were compared by using sensitivities, specificities, predictive values, and Dice coefficient.Results: A total of 312 men (median age, 64 years; interquartile range [IQR], 58-71 years) were evaluated. The training set consisted of 250 men (median age, 64 years; IQR, 58-71 years) and the test set of 62 men (median age, 64 years; IQR, 60-69 years). In the test set, PI-RADS cutoffs greater than or equal to 3 versus cutoffs greater than or equal to 4 on a per-patient basis had sensitivity of 96% (25 of 26) versus 88% (23 of 26) at specificity of 22% (eight of 36) versus 50% (18 of 36). U-Net at probability thresholds of greater than or equal to 0.22 versus greater than or equal to 0.33 had sensitivity of 96% (25 of 26) versus 92% (24 of 26) (both P>.99) with specificity of 31% (11 of 36) versus 47% (17 of 36) (both P>.99), not statistically different from PI-RADS. Dice coefficients were 0.89 for prostate and 0.35 for MRI lesion segmentation. In the test set, coincidence of PI-RADS greater than or equal to 4 with U-Net lesions improved the positive predictive value from 48% (28 of 58) to 67% (24 of 36) for U-Net probability thresholds greater than or equal to 0.33 (P = .01), while the negative predictive value remained unchanged (83% [25 of 30] vs 83%[43 of 52]; P>.99).Conclusion: U-Net trained with T2-weighted and diffusion MRI achieves similar performance to clinical Prostate Imaging Reporting and Data System assessment. (C) RSNA, 2019