Machine Learning for Early Lung Cancer Identification Using Routine Clinical and Laboratory Data

Machine Learning for Early Lung Cancer Identification Using Routine Clinical and Laboratory Data
复制标题

DOI:
10.1164/rccm.202007-2791oc
复制
发表时间:
2021-08-15
影响因子:
24.7
通讯作者:
Shiff, Ron
Shiff, Ron
中科院分区:
医学1区
文献类型:
--
作者:
Gould, Michael K.;Huang, Brian Z.;Shiff, Ron

文献摘要

被引文献

相似文献

理由:大多数肺癌在诊断时已处于晚期。对高危个体进行症状前识别可以促进早期干预并改善长期结果。目标:利用机器学习,开发一种模型,根据常规临床和实验室数据预测未来的肺癌诊断。方法:我们收集了 6,505 名非小细胞肺癌 (NSCLC) 病例患者和 189,597 名同期对照受试者的数据,并将新型机器学习模型的准确性与传统机器学习模型的准确性进行了比较。 经过充分验证的 2012 年前列腺癌、肺癌、结直肠癌和卵巢癌筛查试验风险模型 (mPLCOm2012) 的修改版本,使用受试者工作特征曲线下面积 (AUC)、灵敏度和诊断优势比 (OR) 作为模型性能的衡量标准。测量和主要结果:在测试集中的曾经吸烟者中,机器学习模型比 mPLCOm2012 更准确地识别 临床诊断前 9-12 个月的 NSCLC (P < 0.00001),AUC 为 0.86,诊断 OR 为 12.3,敏感性为 40.1%,预定义特异性为 95%。相比之下,mPLCOm2012 在相同特异性下的 AUC 为 0.79,OR 为 7.4,灵敏度为 27.9%。当应用于符合筛查资格的人群时,机器学习模型比肺癌筛查的标准资格标准更准确,并且比 mPLCOm2012 更准确。影响模型变量包括已知的风险因素和白细胞和血小板计数等新的预测因子。结论:机器学习模型对于 NSCLC 的早期诊断比标准筛查资格标准或 mPLCOm2012 更准确,这表明通过早期检测有助于预防肺癌死亡的潜力。
Rationale: Most lung cancers are diagnosed at an advanced stage. Presymptomatic identification of high-risk individuals can prompt earlier intervention and improve long-term outcomes.Objectives: To develop a model to predict a future diagnosis of lung cancer on the basis of routine clinical and laboratory data by using machine learning.Methods: We assembled data from 6,505 case patients with non-small cell lung cancer (NSCLC) and 189,597 contemporaneous control subjects and compared the accuracy of a novel machine learning model with a modified version of the well-validated 2012 Prostate, Lung, Colorectal and Ovarian Cancer Screening Trial risk model (mPLCOm2012), by using the area under the receiver operating characteristic curve (AUC), sensitivity, and diagnostic odds ratio (OR) as measures of model performance.Measurements and Main Results: Among ever-smokers in the test set, a machine learning model was more accurate than the mPLCOm2012 for identifying NSCLC 9-12 months before clinical diagnosis (P < 0.00001) and demonstrated an AUC of 0.86, a diagnostic OR of 12.3, and a sensitivity of 40.1% at a predefined specificity of 95%. In comparison, the mPLCOm2012 demonstrated an AUC of 0.79, an OR of 7.4, and a sensitivity of 27.9% at the same specificity. The machine learning model was more accurate than standard eligibility criteria for lung cancer screening and more accurate than the mPLCOm2012 when applied to a screening-eligible population. Influential model variables included known risk factors and novel predictors such as white blood cell and platelet counts.Conclusions: A machine learning model was more accurate for early diagnosis of NSCLC than either standard eligibility criteria for screening or the mPLCOm2012, demonstrating the potential to help prevent lung cancer deaths through early detection.