Combining the strengths of radiologists and AI for breast cancer screening: a retrospective analysis.

Combining the strengths of radiologists and AI for breast cancer screening: a retrospective analysis.
复制标题

结合放射科医生和人工智能的优势进行乳腺癌筛查:一项回顾性分析。

DOI:
10.1016/s2589-7500(22)00070-x
复制
发表时间:
2022-07
影响因子:
30.8
通讯作者:
Umutlut, Lale
Umutlut, Lale
中科院分区:
医学1区
文献类型:
--
作者:
Leibig, Christian;Brehmer, Moritz;Bunk, Stefan;Byng, Danaiyn;Pinkert, Katja;Umutlut, Lale

文献摘要

被引文献

相似文献

我们提出了一种将人工智能(AI)整合到乳腺癌筛查途径中的决策推荐方法,该算法根据其量化的不确定性进行预测。具有高确定性的放射学评估是自动完成的,而具有较低确定性的评估是由放射科医生进行的。这个由两部分组成的人工智能系统可以对正常的乳房X光检查进行分类,并提供事后癌症检测,以保持高度的灵敏度。本研究旨在评价该AI系统在作为独立系统或在决策-转诊方法中使用时与原始放射科医生决策相比的灵敏度和特异性性能。我们使用了一个回顾性数据集,包括2007年1月1日至2020年12月31日期间进行的1193197项全视野数字乳腺X射线摄影研究,这些研究来自参与德国国家乳腺癌筛查计划的8个筛查点。我们从六个筛选站点获得了内部测试数据集(1670例筛查发现的癌症和19997例正常乳腺X线摄影检查),和乳腺癌筛查的外部测试数据集(2793例筛查发现癌症,80058例正常检查)来自两个额外的筛选研究中心,以评价AI算法在作为独立系统使用时的灵敏度和特异性性能,与在共识会议之前的屏幕点阅读处的原始个体放射科医师决定相比,在决策-转诊方法内,对AI算法的不同配置进行了评价。为了解释过度采样癌症病例导致的数据集丰富,应用权重以反映筛查计划中研究类型的实际分布。将分类性能评价为正确识别为正常的检查率。在独立AI、放射科医生和决策转诊之间比较了临床相关亚组、筛选中心和器械制造商的敏感性。我们提出了受试者工作特征(ROC)曲线和ROC下面积(AUROC),以评估AI系统在整个工作范围内的性能。与放射科医师的比较和亚组分析是基于临床相关配置的灵敏度和特异性。AI系统在独立模式下的示例性配置实现了84.2%的灵敏度(95% CI 82.4 - 85.8),特异性为89.5%(89.0~89.9),灵敏度为84.6%(83.3~85.9),特异性为91.3%(91.1~91.5),但准确性低于放射科医师的平均水平。相比之下,模拟决策-转诊方法显著提高了放射科医生的敏感性2.6个百分点,特异性1.0个百分点,对应于外部数据集上63.0%的分诊性能;在AI评估的研究子集上,AUROC为0.982(95%CI 0.978 - 0.986),超过了放射科医生的性能。决策-转诊方法也使许多临床相关亚组的敏感性显著增加,包括小病变大小和浸润性癌亚组。在纳入的8个筛选中心和3个器械制造商中,决策-转诊方法的敏感性是一致的。决策-转诊方法利用了放射科医生和人工智能的优势,证明了灵敏度和特异性的提高超过了单个放射科医生和独立的人工智能系统。这种方法有可能提高放射科医生的筛查准确性,适应筛查的要求,并可以允许减少共识会议之前的工作量,而不丢弃放射科医生的一般知识。瓦拉
We propose a decision-referral approach for integrating artificial intelligence (AI) into the breast-cancer screening pathway, whereby the algorithm makes predictions on the basis of its quantification of uncertainty. Algorithmic assessments with high certainty are done automatically, whereas assessments with lower certainty are referred to the radiologist. This two-part AI system can triage normal mammography exams and provide post-hoc cancer detection to maintain a high degree of sensitivity. This study aimed to evaluate the performance of this AI system on sensitivity and specificity when used either as a standalone system or within a decision-referral approach, compared with the original radiologist decision. We used a retrospective dataset consisting of 1 193 197 full-field, digital mammography studies carried out between Jan 1, 2007, and Dec 31, 2020, from eight screening sites participating in the German national breast-cancer screening programme. We derived an internal-test dataset from six screening sites (1670 screen-detected cancers and 19 997 normal mammography exams), and an external-test dataset of breast cancer screening exams (2793 screen-detected cancers and 80 058 normal exams) from two additional screening sites to evaluate the performance of an AI algorithm on sensitivity and specificity when used either as a standalone system or within a decision-referral approach, compared with the original individual radiologist decision at the point-of-screen reading ahead of the consensus conference. Different configurations of the AI algorithm were evaluated. To account for the enrichment of the datasets caused by oversampling cancer cases, weights were applied to reflect the actual distribution of study types in the screening programme. Triaging performance was evaluated as the rate of exams correctly identified as normal. Sensitivity across clinically relevant subgroups, screening sites, and device manufacturers was compared between standalone AI, the radiologist, and decision referral. We present receiver operating characteristic (ROC) curves and area under the ROC (AUROC) to evaluate AI-system performance over its entire operating range. Comparison with radiologists and subgroup analysis was based on sensitivity and specificity at clinically relevant configurations. The exemplary configuration of the AI system in standalone mode achieved a sensitivity of 84·2% (95% CI 82·4–85·8) and a specificity of 89·5% (89·0–89·9) on internal-test data, and a sensitivity of 84·6% (83·3–85·9) and a specificity of 91·3% (91·1–91·5) on external-test data, but was less accurate than the average unaided radiologist. By contrast, the simulated decision-referral approach significantly improved upon radiologist sensitivity by 2·6 percentage points and specificity by 1·0 percentage points, corresponding to a triaging performance at 63·0% on the external dataset; the AUROC was 0·982 (95% CI 0·978–0·986) on the subset of studies assessed by AI, surpassing radiologist performance. The decision-referral approach also yielded significant increases in sensitivity for a number of clinically relevant subgroups, including subgroups of small lesion sizes and invasive carcinomas. Sensitivity of the decision-referral approach was consistent across the eight included screening sites and three device manufacturers. The decision-referral approach leverages the strengths of both the radiologist and AI, demonstrating improvements in sensitivity and specificity surpassing that of the individual radiologist and of the standalone AI system. This approach has the potential to improve the screening accuracy of radiologists, is adaptive to the requirements of screening, and could allow for the reduction of workload ahead of the consensus conference, without discarding the generalised knowledge of radiologists. Vara.