Predicting diabetes mellitus using SMOTE and ensemble machine learning approach: The Henry Ford ExercIse Testing (FIT) project.

Predicting diabetes mellitus using SMOTE and ensemble machine learning approach: The Henry Ford ExercIse Testing (FIT) project.
复制标题

DOI:
10.1371/journal.pone.0179805
复制
发表时间:
2017
期刊:
影响因子:
3.7
通讯作者:
Sakr S
Sakr S
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Alghamdi M;Al-Mallah M;Keteyian S;Brawner C;Ehrman J;Sakr S

文献摘要

参考文献

被引文献

相似文献

机器学习正在成为医学研究领域的一种流行而重要的方法。在这项研究中,我们调查了各种机器学习方法的相对性能,如决策树,朴素贝叶斯,逻辑回归,逻辑模型树和随机森林,用于使用心肺健康的医疗记录预测糖尿病事件。此外,我们应用不同的技术来发现糖尿病的潜在预测因素。这项FIT项目研究使用了1991年至2009年期间在亨利福特健康系统接受临床医生推荐的运动平板负荷试验的32,555例无任何已知冠状动脉疾病或心力衰竭患者的数据,并进行了完整的5年随访。在第五年结束时,其中5 099名患者患上了糖尿病。该数据集包含62个属性,分为四类:人口统计学特征,疾病史,药物使用史和压力测试生命体征。我们开发了一个基于集成的预测模型,使用13个属性,这些属性是根据其临床重要性,多元线性回归和信息增益排名方法选择的。通过合成少数过采样技术(SMOTE)处理所构建模型的不平衡类的负面影响。预测模型分类器的整体性能通过Ensemble机器学习方法使用具有三个决策树(朴素贝叶斯树、随机森林和逻辑模型树)的投票方法来提高,并实现了高预测精度(AUC = 0.92)。该研究显示了使用心肺健康数据预测糖尿病事件的集成和SMOTE方法的潜力。
Machine learning is becoming a popular and important approach in the field of medical research. In this study, we investigate the relative performance of various machine learning methods such as Decision Tree, Naïve Bayes, Logistic Regression, Logistic Model Tree and Random Forests for predicting incident diabetes using medical records of cardiorespiratory fitness. In addition, we apply different techniques to uncover potential predictors of diabetes. This FIT project study used data of 32,555 patients who are free of any known coronary artery disease or heart failure who underwent clinician-referred exercise treadmill stress testing at Henry Ford Health Systems between 1991 and 2009 and had a complete 5-year follow-up. At the completion of the fifth year, 5,099 of those patients have developed diabetes. The dataset contained 62 attributes classified into four categories: demographic characteristics, disease history, medication use history, and stress test vital signs. We developed an Ensembling-based predictive model using 13 attributes that were selected based on their clinical importance, Multiple Linear Regression, and Information Gain Ranking methods. The negative effect of the imbalance class of the constructed model was handled by Synthetic Minority Oversampling Technique (SMOTE). The overall performance of the predictive model classifier was improved by the Ensemble machine learning approach using the Vote method with three Decision Trees (Naïve Bayes Tree, Random Forest, and Logistic Model Tree) and achieved high accuracy of prediction (AUC = 0.92). The study shows the potential of ensembling and SMOTE approaches for predicting incident diabetes using cardiorespiratory fitness data.
DOI: 10.5539/gjhs.v7n5p304
发表时间: 2015-03-18
期刊: Global journal of health science
影响因子: --
作者:
Habibi S;Ahmadi M;Alizadeh S
通讯作者: Alizadeh S
DOI: 10.1136/bmjopen-2012-002457
发表时间: 2013-05-14
期刊: BMJ open
影响因子: 2.9
作者:
Farran B;Channanath AM;Behbehani K;Thanaraj TA
通讯作者: Thanaraj TA
DOI: 10.1093/biomet/70.1.163
发表时间: 1983-01-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
KENT, JT
通讯作者: KENT, JT
DOI: 10.1186/s12859-015-0723-9
发表时间: 2015-09-21
期刊: BMC bioinformatics
影响因子: 3
作者:
Blagus R;Lusa L
通讯作者: Lusa L
DOI: 10.1016/j.csda.2009.04.009
发表时间: 2009-09-01
影响因子: 1.8
作者:
Kim, Ji-Hyun
通讯作者: Kim, Ji-Hyun