Comparison and development of machine learning tools for the prediction of chronic obstructive pulmonary disease in the Chinese population

Comparison and development of machine learning tools for the prediction of chronic obstructive pulmonary disease in the Chinese population
复制标题

DOI:
10.1186/s12967-020-02312-0
复制
发表时间:
2020-03-31
影响因子:
7.4
通讯作者:
Jia, Weihua
Jia, Weihua
中科院分区:
医学2区
文献类型:
--
作者:
Ma, Xia;Wu, Yanping;Jia, Weihua

文献摘要

被引文献

相似文献

背景慢性阻塞性肺疾病(COPD)是一个主要的公共卫生问题和全球范围内的死亡原因。然而,早期COPD通常无法识别和诊断。因此,有必要建立一个预测COPD发展的危险模型。方法采用MassArray技术检测441例COPD患者和192例健康对照者的101个单核苷酸多态性(SNPs)。利用5个临床特征和SNPs建立6个预测模型,并通过混淆矩阵AU-ROC、AU-PRC、灵敏度(召回率)、特异度、准确度、F1评分、MCC、PPV(精确度)和NPV在训练集和测试集上进行评价。对选定的特征进行了排序。结果9个SNPs与COPD的发病有显著相关性。其中,6个SNPs(rs1007052,OR = 1.671,P = 0.010; rs2910164,OR = 1.416,P < 0.037; rs473892,OR = 1.473,P < 0.044; rs161976,OR = 1.594,P < 0.044; rs 159497,OR = 1.445,P < 0.045; rs 9296092,OR = 1.832,P < 0.045)是COPD的危险因素,3个SNPs rs 8192288,OR = 0.593,P < 0.015; rs 20541,OR = 0.669,P < 0.018; rs 12922394,OR = 0.651,P < 0.022)是COPD发展的保护因素。在训练集中,KNN、LR、SVM、DT和XGboost获得的AU-ROC值高于0.82,AU-PRC值高于0.92。在这些模型中,XGboost获得了最高的AU-ROC(0.94),AU-PRC(0.97),准确度(0.91),精确度(0.95),F1得分(0.94),MCC(0.77)和特异性(0.85),而MLP获得了最高的灵敏度(召回)(0.99)和NPV(0.87)。在验证集中,KNN,LR和XGboost分别获得了高于0.80和0.85的AU-ROC和AU-PRC值。KNN具有最高的精度(0.82),KNN和LR获得同样的最高准确度(0.81),KNN和LR具有同样的最高F1得分(0.86)。DT和MLP的灵敏度(回忆)和NPV值分别高于0.94和0.84。在特征重要性分析中,我们发现AQCI、年龄和BMI对模型的预测能力影响最大,而SNP、性别和吸烟则不太重要。结论KNN、LR和XGboost模型显示出良好的整体预测能力,结合临床和SNP特征的机器学习工具适用于预测COPD发展的风险。
Background Chronic obstructive pulmonary disease (COPD) is a major public health problem and cause of mortality worldwide. However, COPD in the early stage is usually not recognized and diagnosed. It is necessary to establish a risk model to predict COPD development. Methods A total of 441 COPD patients and 192 control subjects were recruited, and 101 single-nucleotide polymorphisms (SNPs) were determined using the MassArray assay. With 5 clinical features as well as SNPs, 6 predictive models were established and evaluated in the training set and test set by the confusion matrix AU-ROC, AU-PRC, sensitivity (recall), specificity, accuracy, F1 score, MCC, PPV (precision) and NPV. The selected features were ranked. Results Nine SNPs were significantly associated with COPD. Among them, 6 SNPs (rs1007052, OR = 1.671, P = 0.010; rs2910164, OR = 1.416, P < 0.037; rs473892, OR = 1.473, P < 0.044; rs161976, OR = 1.594, P < 0.044; rs159497, OR = 1.445, P < 0.045; and rs9296092, OR = 1.832, P < 0.045) were risk factors for COPD, while 3 SNPs (rs8192288, OR = 0.593, P < 0.015; rs20541, OR = 0.669, P < 0.018; and rs12922394, OR = 0.651, P < 0.022) were protective factors for COPD development. In the training set, KNN, LR, SVM, DT and XGboost obtained AU-ROC values above 0.82 and AU-PRC values above 0.92. Among these models, XGboost obtained the highest AU-ROC (0.94), AU-PRC (0.97), accuracy (0.91), precision (0.95), F1 score (0.94), MCC (0.77) and specificity (0.85), while MLP obtained the highest sensitivity (recall) (0.99) and NPV (0.87). In the validation set, KNN, LR and XGboost obtained AU-ROC and AU-PRC values above 0.80 and 0.85, respectively. KNN had the highest precision (0.82), both KNN and LR obtained the same highest accuracy (0.81), and KNN and LR had the same highest F1 score (0.86). Both DT and MLP obtained sensitivity (recall) and NPV values above 0.94 and 0.84, respectively. In the feature importance analyses, we identified that AQCI, age, and BMI had the greatest impact on the predictive abilities of the models, while SNPs, sex and smoking were less important. Conclusions The KNN, LR and XGboost models showed excellent overall predictive power, and the use of machine learning tools combining both clinical and SNP features was suitable for predicting the risk of COPD development.