A data-driven approach to predicting diabetes and cardiovascular disease with machine learning

A data-driven approach to predicting diabetes and cardiovascular disease with machine learning
复制标题

DOI:
10.1186/s12911-019-0918-5
复制
发表时间:
2019-11-06
影响因子:
3.5
通讯作者:
Mohanty, Somya D.
Mohanty, Somya D.
中科院分区:
医学3区
文献类型:
--
作者:
Dinh, An;Miertschin, Stacey;Mohanty, Somya D.

文献摘要

被引文献

相似文献

背景在美国,糖尿病和心血管疾病是两种主要的死亡原因。在患者中识别和预测这些疾病是阻止其进展的第一步。我们使用调查数据(和实验室结果)评估机器学习模型在检测高危患者方面的能力,并在数据中确定在患者中导致这些疾病的关键变量。方法我们的研究探索数据驱动的方法,利用有监督的机器学习模型来识别患有此类疾病的患者。使用国家健康和营养检查调查(NHANES)数据集,我们对数据中所有可用的特征变量进行了详尽的搜索,以开发用于心血管、糖尿病前期和糖尿病检测的模型。使用不同的时间框架和数据特征集(基于实验室数据),对多种机器学习模型(Logistic回归、支持向量机、随机森林和梯度提升)的分类性能进行了评估。然后,将这些模型组合起来开发一个加权集成模型,该模型能够利用不同模型的性能来提高检测精度。基于树的模型的信息增益被用来识别患者数据中的关键变量,这些变量有助于通过数据学习模型检测每种疾病类别中的高危患者。结果开发的心血管疾病集成模型(基于131个变量)在没有实验室结果的情况下获得了83.1%的面积欠受试者操作特征(AU-ROC)分数,与实验室结果的准确率为83.9%。在糖尿病分类中(基于123个变量),极限梯度增强(XGBoost)模型获得了86.2%(没有实验室数据)和95.7%(有实验室数据)的AU-ROC得分。对于糖尿病前期患者,集成模型的AU-ROC得分最高,为73.7%(没有实验室数据),而对于基于实验室的数据,XGBoost表现最好,为84.4%。糖尿病患者的前五个预测因素是1)腰围、2)年龄、3)自我报告体重、4)腿长和5)钠摄入量。对于心血管疾病,模型确定了1)年龄,2)收缩压,3)自我报告体重,4)胸痛的发生,以及5)舒张压是关键因素。结论基于调查问卷的机器学习模型可以为糖尿病和心血管疾病高危患者提供一种自动识别机制。我们还确定了预测的关键贡献者,可以进一步探讨它们对电子健康记录的影响。
Background Diabetes and cardiovascular disease are two of the main causes of death in the United States. Identifying and predicting these diseases in patients is the first step towards stopping their progression. We evaluate the capabilities of machine learning models in detecting at-risk patients using survey data (and laboratory results), and identify key variables within the data contributing to these diseases among the patients. Methods Our research explores data-driven approaches which utilize supervised machine learning models to identify patients with such diseases. Using the National Health and Nutrition Examination Survey (NHANES) dataset, we conduct an exhaustive search of all available feature variables within the data to develop models for cardiovascular, prediabetes, and diabetes detection. Using different time-frames and feature sets for the data (based on laboratory data), multiple machine learning models (logistic regression, support vector machines, random forest, and gradient boosting) were evaluated on their classification performance. The models were then combined to develop a weighted ensemble model, capable of leveraging the performance of the disparate models to improve detection accuracy. Information gain of tree-based models was used to identify the key variables within the patient data that contributed to the detection of at-risk patients in each of the diseases classes by the data-learned models. Results The developed ensemble model for cardiovascular disease (based on 131 variables) achieved an Area Under - Receiver Operating Characteristics (AU-ROC) score of 83.1% using no laboratory results, and 83.9% accuracy with laboratory results. In diabetes classification (based on 123 variables), eXtreme Gradient Boost (XGBoost) model achieved an AU-ROC score of 86.2% (without laboratory data) and 95.7% (with laboratory data). For pre-diabetic patients, the ensemble model had the top AU-ROC score of 73.7% (without laboratory data), and for laboratory based data XGBoost performed the best at 84.4%. Top five predictors in diabetes patients were 1) waist size, 2) age, 3) self-reported weight, 4) leg length, and 5) sodium intake. For cardiovascular diseases the models identified 1) age, 2) systolic blood pressure, 3) self-reported weight, 4) occurrence of chest pain, and 5) diastolic blood pressure as key contributors. Conclusion We conclude machine learned models based on survey questionnaire can provide an automated identification mechanism for patients at risk of diabetes and cardiovascular diseases. We also identify key contributors to the prediction, which can be further explored for their implications on electronic health records.