Stand-Alone Artificial Intelligence for Breast Cancer Detection in Mammography: Comparison With 101 Radiologists

Stand-Alone Artificial Intelligence for Breast Cancer Detection in Mammography: Comparison With 101 Radiologists
复制标题

DOI:
10.1093/jnci/djy222
复制
发表时间:
2019-09-01
影响因子:
10.3
通讯作者:
Sechopoulos, Ioannis
Sechopoulos, Ioannis
中科院分区:
医学1区
文献类型:
--
作者:
Rodriguez-Ruiz, Alejandro;Lang, Kristina;Sechopoulos, Ioannis

文献摘要

被引文献

相似文献

背景资料:人工智能(AI)系统在评估数字乳腺X射线摄影(DM)时的放射学水平将提高乳腺癌筛查的准确性和效率。我们的目的是比较独立的性能的AI系统的放射科医生在检测乳腺癌的DM。方法:9个多读者,多病例研究数据集,以前用于不同的研究目的,在7个国家收集。每个数据集包括使用来自四个不同供应商的系统采集的DM检查、每次检查的多名放射科医生评估以及通过组织病理学分析或随访验证的地面实况,总共产生2652次检查(653次恶性)和101名放射科医生的解释(28296次独立解释)。一个人工智能系统分析了这些检查结果,得出了1到10之间的癌症怀疑水平。采用非劣效性零假设(0.05)比较放射科医师与人工智能系统的检测性能。结果:人工智能系统的检测性能在统计学上不劣于101名放射科医师的平均值。AI系统的ROC曲线下面积为0.840(95%置信区间[CI] = 0.820至0.860),放射科医师的平均值为0.814(95% CI = 0.787至0.841)(差异95% CI = -0.003至0.055)。AI系统的AUC高于61.4%of the radiologist.Conclusions:评估的AI系统实现了癌症检测的准确性相比,平均乳腺放射科医生在这种回顾性设置。虽然有希望,这样的系统在筛选设置的性能和影响需要进一步调查。
Background: Artificial intelligence (AI) systems performing at radiologist-like levels in the evaluation of digital mammography (DM) would improve breast cancer screening accuracy and efficiency. We aimed to compare the stand-alone performance of an AI system to that of radiologists in detecting breast cancer in DM.Methods: Nine multi-reader, multi-case study datasets previously used for different research purposes in seven countries were collected. Each dataset consisted of DM exams acquired with systems from four different vendors, multiple radiologists' assessments per exam, and ground truth verified by histopathological analysis or follow-up, yielding a total of 2652 exams (653 malignant) and interpretations by 101 radiologists (28 296 independent interpretations). An AI system analyzed these exams yielding a level of suspicion of cancer present between 1 and 10. The detection performance between the radiologists and the AI system was compared using a noninferiority null hypothesis at a margin of 0.05.Results: The performance of the AI system was statistically noninferior to that of the average of the 101 radiologists. The AI system had a 0.840 (95% confidence interval [CI] = 0.820 to 0.860) area under the ROC curve and the average of the radiologists was 0.814 (95% CI = 0.787 to 0.841) (difference 95% CI = -0.003 to 0.055). The AI system had an AUC higher than 61.4% of the radiologists.Conclusions: The evaluated AI system achieved a cancer detection accuracy comparable to an average breast radiologist in this retrospective setting. Although promising, the performance and impact of such a system in a screening setting needs further investigation.