External validation of a publicly available computer assisted diagnostic tool for mammographic mass lesions with two high prevalence research datasets.

External validation of a publicly available computer assisted diagnostic tool for mammographic mass lesions with two high prevalence research datasets.
复制标题

DOI:
10.1118/1.4927260
复制
发表时间:
2015-08
期刊:
影响因子:
3.8
通讯作者:
M. Benndorf;E. Burnside;Christoph Herda;M. Langer;E. Kotter
M. Benndorf;E. Burnside;Christoph Herda;M. Langer;E. Kotter
中科院分区:
医学3区
文献类型:
--
作者:
M. Benndorf;E. Burnside;Christoph Herda;M. Langer;E. Kotter

文献摘要

被引文献

相似文献

目的 用高度标准化的术语描述乳房 X 光检查中检测到的病变:乳房成像报告和数据系统 (BI-RADS) 词典。到目前为止,还没有经过验证的语义计算机辅助分类算法可以将词典中的形态描述符组合交互式链接到恶性肿瘤的概率风险估计。因此,作者的目标是对乳房 X 光肿块诊断 (MMassDx) 算法进行外部验证。像 MMassDx 这样的分类算法必须在各种临床环境和未用于生成算法的数据集中表现良好,才能最终被临床常规所接受。方法 MMassDx 算法使用朴素贝叶斯网络,并根据两组不同的变量计算测试后的恶性肿瘤概率:(a) BI-RADS 描述符和年龄(“描述符模型”)和 (b) BI-RADS 描述符、年龄和 BI-RADS 评估类别(“包容性模型”)。作者使用两个大型公开可用的乳腺 X 线摄影肿块病变数据集评估 MMassDx(描述符)和 MMassDx(含)模型:用于筛查乳腺 X 线摄影的数字数据库 (DDSM) 数据集,其中包含来自同一检查的两个子集 - 中外侧斜 (MLO) 视图和头尾 (CC) 视图数据集 - 以及乳房 X 线摄影肿块 (MM) 数据集。 DDSM 包含 1220 个肿块病变,MM 数据集包含 961 个肿块病变。作者使用接受者操作特征曲线 (AUC) 下的面积评估判别性能,并使用 DeLong 方法将其与单独的 BI-RADS 评估类别(即临床性能)进行比较。作者还使用校准曲线评估指定的概率风险估计是否反映了病变的真实恶性肿瘤风险。结果 作者证明 MMassDx 算法表现出良好的判别性能。 DDSM 数据中 MMassDx(描述符)模型的 AUC 为 0.876/0.895(MLO/CC 视图),DDSM 数据中 MMassDx(包含)模型的 AUC 为 0.891/0.900(MLO/CC 视图)。 MM 数据中 MMassDx(描述符)模型的 AUC 为 0.862,MM 数据中 MMassDx(含)模型的 AUC 为 0.900。在所有情况下,MMassDx 的表现均显着优于临床表现,每种情况 P < 0.05。作者进一步证明,MMassDx 算法系统地低估了 DDSM 和 MM 数据集中的恶性肿瘤风险,特别是当分配的恶性肿瘤概率较低时。结论 作者的结果表明,在两个独立的验证数据集上进行测试时,MMassDx 算法具有良好的区分性能,但校准精度较低。在未来的临床人群中改进校准和测试将是将这些算法转化为临床的重要步骤。
PURPOSE Lesions detected at mammography are described with a highly standardized terminology: the breast imaging-reporting and data system (BI-RADS) lexicon. Up to now, no validated semantic computer assisted classification algorithm exists to interactively link combinations of morphological descriptors from the lexicon to a probabilistic risk estimate of malignancy. The authors therefore aim at the external validation of the mammographic mass diagnosis (MMassDx) algorithm. A classification algorithm like MMassDx must perform well in a variety of clinical circumstances and in datasets that were not used to generate the algorithm in order to ultimately become accepted in clinical routine. METHODS The MMassDx algorithm uses a naïve Bayes network and calculates post-test probabilities of malignancy based on two distinct sets of variables, (a) BI-RADS descriptors and age ("descriptor model") and (b) BI-RADS descriptors, age, and BI-RADS assessment categories ("inclusive model"). The authors evaluate both the MMassDx (descriptor) and MMassDx (inclusive) models using two large publicly available datasets of mammographic mass lesions: the digital database for screening mammography (DDSM) dataset, which contains two subsets from the same examinations-a medio-lateral oblique (MLO) view and cranio-caudal (CC) view dataset-and the mammographic mass (MM) dataset. The DDSM contains 1220 mass lesions and the MM dataset contains 961 mass lesions. The authors evaluate discriminative performance using area under the receiver-operating-characteristic curve (AUC) and compare this to the BI-RADS assessment categories alone (i.e., the clinical performance) using the DeLong method. The authors also evaluate whether assigned probabilistic risk estimates reflect the lesions' true risk of malignancy using calibration curves. RESULTS The authors demonstrate that the MMassDx algorithms show good discriminatory performance. AUC for the MMassDx (descriptor) model in the DDSM data is 0.876/0.895 (MLO/CC view) and AUC for the MMassDx (inclusive) model in the DDSM data is 0.891/0.900 (MLO/CC view). AUC for the MMassDx (descriptor) model in the MM data is 0.862 and AUC for the MMassDx (inclusive) model in the MM data is 0.900. In all scenarios, MMassDx performs significantly better than clinical performance, P < 0.05 each. The authors furthermore demonstrate that the MMassDx algorithm systematically underestimates the risk of malignancy in the DDSM and MM datasets, especially when low probabilities of malignancy are assigned. CONCLUSIONS The authors' results reveal that the MMassDx algorithms have good discriminatory performance but less accurate calibration when tested on two independent validation datasets. Improvement in calibration and testing in a prospective clinical population will be important steps in the pursuit of translation of these algorithms to the clinic.