Artificial Intelligence and Data Mining to Assess Lung Cancer Risk: Challenges and Opportunities
Artificial Intelligence and Data Mining to Assess Lung Cancer Risk: Challenges and Opportunities
复制标题
人工智能和数据挖掘评估肺癌风险:挑战和机遇
DOI:
10.7326/m20-5673
复制
发表时间:
2020
影响因子:
39.2
通讯作者:
P. Pinsky
中科院分区:
文献类型:
--
作者:
P. Pinsky
Challenges and Opportunities T article by Lu and colleagues describes a creative use of artificial intelligence (AI) and stored patient data to assess risk for lung cancer, with the aim of identifying persons potentially eligible for screening (1). This research brings up several points of interest. Most germane to the article are how to assess eligibility for low-dose computed tomography (LDCT) lung cancer screening and how to understand the nature of AI in the context of lung cancer detection and prediction. However, viewing the research in its wider context spotlights many issues associated with the use of AI and, more generally, data mining of electronic health records (EHRs) to improve patient care. The CXR-LC model developed by the study authors, a deep-learning application using stored chest radiographs (CXRs) and basic demographic data (age, sex, and current smoking status), showed better prediction of future incident lung cancer than the current standard LDCT eligibility criteria, which are based only on age, pack-years, and years since quitting, and similar predictive ability to the PLCO (Prostate, Lung, Colorectal, and Ovarian) Cancer Screening Trial Model 2012 (PLCOM2012), a risk model based on detailed smoking history and demographic characteristics (1). The U.S. Preventive Services Task Force is currently updating its recommendation for lung cancer screening (2, 3). The draft recommendation states that there was insufficient evidence to “assess whether or not risk prediction model–based screening would improve outcomes” (3). An argument against using current risk models (for example, PLCOM2012) is that they incorporate age as a risk factor and thus skew eligibility toward older persons. Although older persons are generally at higher risk for lung cancer, they also have fewer potential life-years saved from screening, so using risk-based criteria may increase the number of lives saved but not necessarily the years of life saved (3). In the CXR-LC model, mean age increased with increasing risk score percentiles, similarly to the PLCOM2012, implying that this same issue applies. Another concern of the task force was that “use of a more complex, computer-based risk prediction model to determine eligibility could impose a barrier to wider implementation” of lung cancer screening, especially given the very low uptake of such screening (3). This concern would presumably be greater for an AI-based prediction tool than a standard risk model. Lu and colleagues present the CXR-LC model primarily as a tool for identifying persons who are potentially eligible for screening by the standard criteria, in settings where EHRs are insufficient to make that determination, with the goal of increasing uptake. Indeed, studies show that more than half the time EHRs lack sufficient smoking history data to determine pack-years and years since quitting, even under managed care (4). However, in the future, it is likely that there will generally be a greater push to use models, including AI tools, to assess risk and determine eligibility for interventions. What is this AI tool actually measuring? An important distinction is between risk prediction and early detection. Presumably, the CXR-LC model is doing both, given that it was trained to detect cancer 0 to 12 years after the CXR. In the highand very-high-risk groups, about one quarter of diagnoses were within 2 years of the CXR, meaning these likely were already (invasive) cancers. To the extent that the model is finding actual cancer, given that CXR screening has been shown to not have a lung cancer mortality benefit, referral to LDCT screening may not be beneficial (5). For risk prediction, the algorithm could be picking up several signals. First, it could be noting the general effects of a heavy smoking history. Second, it could be recognizing a field effect; tobacco smoke induces molecular alterations representing an etiologic field of injury associated with lung carcinogenesis (6). Finally, it could be detecting lung cancer precursors, such as atypical adenomatous hyperplasia; however, because CXRs have poor sensitivity for nodules less than 20 mm, it is unlikely that the model is directly detecting precursors (7). Combining the CXR-LC model with PLCOM2012 added limited predictive ability; therefore, although the metrics are substantially correlated, the CXR-LC model is picking up some different features than PLCOM2012, possibly relating specifically to the carcinogenesis pathway, over and above the general smoking history burden. In the end, the CXR-LC model, like many AI applications, is essentially a black box model. Increasingly, explainability has become an important issue in the medical AI field (8). When a model is not explainable, patients and physicians may lack confidence in its predictions. More generally, in assessing the use of data mining of patient EHRs (including images) for risk prediction, several important issues arise in addition to explainability. First, there is the level of (increased) risk identified and the level of potential mitigation of that risk through interventions. For lung cancer screening, even among eligible persons, the risk for lung cancer death is not that high (about 2% within 6 years) (9). Further, the mitigation of that risk is modest, a 15% to 20% reduction with several rounds of screening. In addition, there are harms of LDCT screening, including false positives and unnecessary invasive procedures (2). These considerations address the magnitude of potential benefit of using an EHR-based prediction tool.