An explainable artificial intelligence framework for risk prediction of COPD in smokers.

An explainable artificial intelligence framework for risk prediction of COPD in smokers.
复制标题

DOI:
10.1186/s12889-023-17011-w
复制
发表时间:
2023-11-06
期刊:
影响因子:
4.5
通讯作者:
--
中科院分区:
医学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

由于与慢性阻塞性肺疾病(COPD)相关的早期体征的不明显性质,个体通常无法识别,导致及时预防和治疗的次优机会。本研究的目的是创建一个可解释的人工智能框架,结合数据预处理方法,机器学习方法和模型可解释性方法,以识别吸烟人群中COPD的高风险人群,并对模型预测提供合理的解释。资料包括问卷调查资料、支气管扩张前后的体格检查资料和肺功能检查结果。首先,采用混合数据析因分析(FAMD)、Boruta和NRSBoundary-SMOTE回归方法解决数据缺失、高维和类别不平衡问题。然后,七个分类模型(CatBoost,NGBoost,XGBoost,LightGBM,随机森林,SVM和逻辑回归)被应用于建模的风险水平,并使用Shapley加法解释(SHAP)方法和偏相关图(PDP)解释最佳机器学习(ML)模型的决策。在吸烟人群中,年龄和其他14个变量是预测COPD的重要因素。CatBoost、随机森林和逻辑回归模型在不平衡数据集中表现相当好。当使用复合指标(AUC,F1评分和G均值)作为模型比较标准时,CatBoost与NRSBoundary-SMOTE在平衡数据集中具有最佳分类性能。年龄、COPD评估试验(CAT)评分、总年收入、体重指数(BMI)、收缩压(SBP)、舒张压(DBP)、窒息、呼吸系统疾病、中心性肥胖、使用污染性燃料取暖、地区、使用污染性燃料做饭和喘息是吸烟人群中预测COPD的重要因素。该研究结合了特征筛选方法、不平衡数据处理方法和先进的机器学习方法,以实现吸烟人群中COPD风险组的早期识别。采用SHAP和PDP识别吸烟人群中的COPD危险因素,旨在为有针对性的筛查策略和吸烟人群自我管理策略提供理论支持。在线版本包含补充材料,可通过10.1186/s12889-023-17011-w获得。
Since the inconspicuous nature of early signs associated with Chronic Obstructive Pulmonary Disease (COPD), individuals often remain unidentified, leading to suboptimal opportunities for timely prevention and treatment. The purpose of this study was to create an explainable artificial intelligence framework combining data preprocessing methods, machine learning methods, and model interpretability methods to identify people at high risk of COPD in the smoking population and to provide a reasonable interpretation of model predictions. The data comprised questionnaire information, physical examination data and results of pulmonary function tests before and after bronchodilatation. First, the factorial analysis for mixed data (FAMD), Boruta and NRSBoundary-SMOTE resampling methods were used to solve the missing data, high dimensionality and category imbalance problems. Then, seven classification models (CatBoost, NGBoost, XGBoost, LightGBM, random forest, SVM and logistic regression) were applied to model the risk level, and the best machine learning (ML) model’s decisions were explained using the Shapley additive explanations (SHAP) method and partial dependence plot (PDP). In the smoking population, age and 14 other variables were significant factors for predicting COPD. The CatBoost, random forest, and logistic regression models performed reasonably well in unbalanced datasets. CatBoost with NRSBoundary-SMOTE had the best classification performance in balanced datasets when composite indicators (the AUC, F1-score, and G-mean) were used as model comparison criteria. Age, COPD Assessment Test (CAT) score, gross annual income, body mass index (BMI), systolic blood pressure (SBP), diastolic blood pressure (DBP), anhelation, respiratory disease, central obesity, use of polluting fuel for household heating, region, use of polluting fuel for household cooking, and wheezing were important factors for predicting COPD in the smoking population. This study combined feature screening methods, unbalanced data processing methods, and advanced machine learning methods to enable early identification of COPD risk groups in the smoking population. COPD risk factors in the smoking population were identified using SHAP and PDP, with the goal of providing theoretical support for targeted screening strategies and smoking population self-management strategies. The online version contains supplementary material available at 10.1186/s12889-023-17011-w.
DOI: 10.7150/ijms.58191
发表时间: 2021
影响因子: 3.6
作者:
Feng Y;Wang Y;Zeng C;Mao H
通讯作者: Mao H
DOI: 10.1109/32.544352
发表时间: 1996-10-01
影响因子: 7.4
作者:
Basili, VR;Briand, LC;Melo, WL
通讯作者: Melo, WL
DOI: 10.1016/s0031-3203(02)00257-1
发表时间: 2003-03-01
影响因子: 8
作者:
Barandela, R;Sánchez, JS;Rangel, E
通讯作者: Rangel, E
一种基于邻域粗糙集模型的边界过采样新算法:NRSBoundary-SMOTE
DOI: 10.1155/2013/694809
发表时间: 2013-01-01
影响因子: --
作者:
Hu, Feng;Li, Hang
通讯作者: Li, Hang
DOI: 10.3760/cma.j.issn.0254-6450.2018.05.002
发表时间: 2018-05-10
影响因子: --
作者:
Fang, L W;Bao, H L;Wang, L H
通讯作者: Wang, L H