A combined strategy of feature selection and machine learning to identify predictors of prediabetes

A combined strategy of feature selection and machine learning to identify predictors of prediabetes
复制标题

DOI:
10.1093/jamia/ocz204
复制
发表时间:
2020-03-01
影响因子:
6.4
通讯作者:
Demmer, Ryan T.
Demmer, Ryan T.
中科院分区:
管理学2区
文献类型:
--
作者:
De Silva, Kushan;Jonsson, Daniel;Demmer, Ryan T.

文献摘要

被引文献

相似文献

目的:利用特征选择和机器学习在美国人口中具有全国代表性的样本中识别糖尿病前期的预测因子。材料与方法:我们分析了2013-2014年全国健康与营养检查调查n = 6346名男性和女性。前驱糖尿病是根据美国糖尿病协会指南定义的。样本被随机分为训练集(n = 3174)和内部验证集(n = 3172)。特征选择算法在包含156个预先选择的暴露变量的训练数据上运行。在使用4种重采样方法构建的原始和重采样训练数据集中,对46个暴露变量应用了4种机器学习算法。预测模型采用2011-2012年全国健康与营养检查调查的内部验证数据(n = 3172)和外部验证数据(n = 3000)进行检验。采用受试者工作特征曲线下面积(AUROC)评价模型性能。预测因子在逻辑模型中采用比值比评估,在其他模型中采用变量重要性评估。疾病控制中心(CDC)的糖尿病前期筛查工具是比较模型性能的基准。结果:糖尿病前期患病率为23.43%。CDC糖尿病前期筛查工具的AUROC为64.40%。7个最优(>= 70% AUROC)模型确定了25个预测因子,包括4个潜在的新关联;逻辑模型和其他非线性/集成模型各占20个,后者仅占5个。所有优化模型都优于疾病预防控制中心糖尿病前期筛查工具(P
Objective: To identify predictors of prediabetes using feature selection and machine learning on a nationally representative sample of the US population.Materials and Methods: We analyzed n = 6346 men and women enrolled in the National Health and Nutrition Examination Survey 2013-2014. Prediabetes was defined using American Diabetes Association guidelines. The sample was randomly partitioned to training (n = 3174) and internal validation (n = 3172) sets. Feature selection algorithms were run on training data containing 156 preselected exposure variables. Four machine learning algorithms were applied on 46 exposure variables in original and resampled training datasets built using 4 resampling methods. Predictive models were tested on internal validation data (n = 3172) and external validation data (n = 3000) prepared from National Health and Nutrition Examination Survey 2011-2012. Model performance was evaluated using area under the receiver operating characteristic curve (AUROC). Predictors were assessed by odds ratios in logistic models and variable importance in others. The Centers for Disease Control (CDC) prediabetes screening tool was the benchmark to compare model performance.Results: Prediabetes prevalence was 23.43%. The CDC prediabetes screening tool produced 64.40% AUROC. Seven optimal (>= 70% AUROC) models identified 25 predictors including 4 potentially novel associations; 20 by both logistic and other nonlinear/ensemble models and 5 solely by the latter. All optimal models outperformed the CDC prediabetes screening tool (P