Artificial Intelligence and Data Mining to Assess Lung Cancer Risk: Challenges and Opportunities

Artificial Intelligence and Data Mining to Assess Lung Cancer Risk: Challenges and Opportunities
复制标题

人工智能和数据挖掘评估肺癌风险:挑战和机遇

DOI:
10.7326/m20-5673
复制
发表时间:
2020
影响因子:
39.2
通讯作者:
P. Pinsky
P. Pinsky
中科院分区:
医学1区
文献类型:
--
作者:
P. Pinsky

文献摘要

被引文献

相似文献

Lu及其同事的挑战和机遇T文章描述了创造性地使用人工智能(AI)和存储的患者数据来评估肺癌风险,目的是识别可能有资格接受筛查的人(1)。这项研究提出了几个有趣的观点。与这篇文章关系最密切的是如何评估低剂量计算机断层扫描(LDCT)肺癌筛查的资格,以及如何在肺癌检测和预测的背景下理解AI的性质。然而,在更广泛的背景下观察这项研究,会发现许多与使用人工智能相关的问题,更普遍的是,电子健康记录(EHR)的数据挖掘,以改善患者护理。研究作者开发的CXR-LC模型,一种使用存储的胸部X光片(CXR)和基本人口统计数据的深度学习应用程序(年龄、性别和当前吸烟状态),显示出比当前标准LDCT资格标准更好的预测未来肺癌事件的能力,该标准仅基于年龄、吸烟包年数和戒烟年数,并且与PLCO相似的预测能力。(前列腺,肺,结直肠和卵巢)癌症筛查试验模型2012(PLCOM 2012),基于详细吸烟史和人口统计学特征的风险模型(1)。美国预防服务工作组目前正在更新其对肺癌筛查的建议(2,3)。该建议草案指出,没有足够的证据来“评估基于风险预测模型的筛查是否会改善结果”(3)。反对使用当前风险模型(例如PLCOM 2012)的一个论点是,它们将年龄作为一个风险因素,从而使老年人的资格倾斜。虽然老年人患肺癌的风险通常较高,但他们从筛查中挽救的潜在生命年数也较少,因此使用基于风险的标准可能会增加挽救的生命数,但不一定是挽救的生命年数。在CXR-LC模型中,与PLCOM 2012相似,平均年龄随着风险评分的增加而增加,这意味着同样的问题也适用。工作组的另一个担忧是,“使用更复杂的基于计算机的风险预测模型来确定资格可能会对肺癌筛查的更广泛实施造成障碍”,特别是考虑到这种筛查的使用率非常低。对于基于人工智能的预测工具来说,这种担忧可能比标准风险模型更大。Lu和他的同事提出CXR-LC模型主要是作为一种工具,用于在EHR不足以做出这一决定的情况下,识别有可能符合标准标准筛查条件的人,目的是增加摄入量。事实上,研究表明,超过一半的时间EHR缺乏足够的吸烟史数据来确定戒烟后的包年数和年数,即使在管理式护理下也是如此(4)。然而,在未来,很可能会有更大的推动力使用模型,包括人工智能工具,以评估风险和确定干预的资格。这个AI工具实际上在测量什么?一个重要的区别是风险预测和早期发现。据推测,CXR-LC模型正在做这两件事,因为它在CXR后0到12年被训练来检测癌症。在高危和极高危人群中,大约四分之一的诊断是在CXR的2年内,这意味着这些可能已经是(侵袭性)癌症。在某种程度上,该模型正在发现实际的癌症,鉴于CXR筛查已被证明没有肺癌死亡率的好处,转介到LDCT筛查可能是没有好处的(5)。对于风险预测,该算法可以拾取多个信号。首先,它可能是注意到大量吸烟史的一般影响。其次,它可能是认识到一个领域的影响;烟草烟雾诱导分子的变化,代表了与肺癌相关的损伤病因领域(6)。最后,它可以检测肺癌前体,如非典型腺瘤样增生;然而,由于CXR对小于20 mm的结节的敏感性较差,因此该模型不太可能直接检测前体(7)。将CXR-LC模型与PLCOM 2012相结合增加了有限的预测能力;因此,尽管指标基本上相关,但CXR-LC模型提取了一些与PLCOM 2012不同的特征,可能与致癌途径特别相关,超过了一般吸烟史负担。最终,CXR-LC模型与许多AI应用一样,本质上是一个黑盒模型。可解释性越来越成为医疗AI领域的一个重要问题(8)。当模型无法解释时,患者和医生可能对其预测缺乏信心。更一般地说,在评估使用患者EHR(包括图像)的数据挖掘进行风险预测时,除了可解释性之外,还出现了几个重要问题。第一,已查明的(增加的)风险程度以及通过干预措施减轻风险的可能程度。对于肺癌筛查,即使在符合条件的人群中,肺癌死亡的风险也不高(6年内约为2%)(9)。此外,这种风险的缓解是适度的,通过几轮筛查减少了15%至20%。此外,LDCT筛查也有危害,包括假阳性和不必要的侵入性操作(2)。这些考虑解决了使用基于EHR的预测工具的潜在益处的大小。
Challenges and Opportunities T article by Lu and colleagues describes a creative use of artificial intelligence (AI) and stored patient data to assess risk for lung cancer, with the aim of identifying persons potentially eligible for screening (1). This research brings up several points of interest. Most germane to the article are how to assess eligibility for low-dose computed tomography (LDCT) lung cancer screening and how to understand the nature of AI in the context of lung cancer detection and prediction. However, viewing the research in its wider context spotlights many issues associated with the use of AI and, more generally, data mining of electronic health records (EHRs) to improve patient care. The CXR-LC model developed by the study authors, a deep-learning application using stored chest radiographs (CXRs) and basic demographic data (age, sex, and current smoking status), showed better prediction of future incident lung cancer than the current standard LDCT eligibility criteria, which are based only on age, pack-years, and years since quitting, and similar predictive ability to the PLCO (Prostate, Lung, Colorectal, and Ovarian) Cancer Screening Trial Model 2012 (PLCOM2012), a risk model based on detailed smoking history and demographic characteristics (1). The U.S. Preventive Services Task Force is currently updating its recommendation for lung cancer screening (2, 3). The draft recommendation states that there was insufficient evidence to “assess whether or not risk prediction model–based screening would improve outcomes” (3). An argument against using current risk models (for example, PLCOM2012) is that they incorporate age as a risk factor and thus skew eligibility toward older persons. Although older persons are generally at higher risk for lung cancer, they also have fewer potential life-years saved from screening, so using risk-based criteria may increase the number of lives saved but not necessarily the years of life saved (3). In the CXR-LC model, mean age increased with increasing risk score percentiles, similarly to the PLCOM2012, implying that this same issue applies. Another concern of the task force was that “use of a more complex, computer-based risk prediction model to determine eligibility could impose a barrier to wider implementation” of lung cancer screening, especially given the very low uptake of such screening (3). This concern would presumably be greater for an AI-based prediction tool than a standard risk model. Lu and colleagues present the CXR-LC model primarily as a tool for identifying persons who are potentially eligible for screening by the standard criteria, in settings where EHRs are insufficient to make that determination, with the goal of increasing uptake. Indeed, studies show that more than half the time EHRs lack sufficient smoking history data to determine pack-years and years since quitting, even under managed care (4). However, in the future, it is likely that there will generally be a greater push to use models, including AI tools, to assess risk and determine eligibility for interventions. What is this AI tool actually measuring? An important distinction is between risk prediction and early detection. Presumably, the CXR-LC model is doing both, given that it was trained to detect cancer 0 to 12 years after the CXR. In the highand very-high-risk groups, about one quarter of diagnoses were within 2 years of the CXR, meaning these likely were already (invasive) cancers. To the extent that the model is finding actual cancer, given that CXR screening has been shown to not have a lung cancer mortality benefit, referral to LDCT screening may not be beneficial (5). For risk prediction, the algorithm could be picking up several signals. First, it could be noting the general effects of a heavy smoking history. Second, it could be recognizing a field effect; tobacco smoke induces molecular alterations representing an etiologic field of injury associated with lung carcinogenesis (6). Finally, it could be detecting lung cancer precursors, such as atypical adenomatous hyperplasia; however, because CXRs have poor sensitivity for nodules less than 20 mm, it is unlikely that the model is directly detecting precursors (7). Combining the CXR-LC model with PLCOM2012 added limited predictive ability; therefore, although the metrics are substantially correlated, the CXR-LC model is picking up some different features than PLCOM2012, possibly relating specifically to the carcinogenesis pathway, over and above the general smoking history burden. In the end, the CXR-LC model, like many AI applications, is essentially a black box model. Increasingly, explainability has become an important issue in the medical AI field (8). When a model is not explainable, patients and physicians may lack confidence in its predictions. More generally, in assessing the use of data mining of patient EHRs (including images) for risk prediction, several important issues arise in addition to explainability. First, there is the level of (increased) risk identified and the level of potential mitigation of that risk through interventions. For lung cancer screening, even among eligible persons, the risk for lung cancer death is not that high (about 2% within 6 years) (9). Further, the mitigation of that risk is modest, a 15% to 20% reduction with several rounds of screening. In addition, there are harms of LDCT screening, including false positives and unnecessary invasive procedures (2). These considerations address the magnitude of potential benefit of using an EHR-based prediction tool.