Type 2 Diabetes Mellitus Screening and Risk Factors Using Decision Tree: Results of Data Mining.

Type 2 Diabetes Mellitus Screening and Risk Factors Using Decision Tree: Results of Data Mining.
复制标题

DOI:
10.5539/gjhs.v7n5p304
复制
发表时间:
2015-03-18
期刊:
Global journal of health science
影响因子:
--
通讯作者:
Alizadeh S
Alizadeh S
中科院分区:
其他
文献类型:
--
作者:
Habibi S;Ahmadi M;Alizadeh S

文献摘要

被引文献

相似文献

这项研究的目的是检验使用与2型糖尿病危险因素相关的特征的预测模型。这些数据是从伊朗大不里士糖尿病控制系统的数据库中获得的。这些数据包括2009至2011年间所有接受糖尿病筛查的人。被认为是“输入”的特征是:年龄、性别、收缩和舒张压、糖尿病家族史和体重指数(BMI)。此外,我们还使用了诊断法作为“类”。我们在WEKA(3.6.10版)软件中应用了“决策树”技术和“J48”算法来开发模型。经过数据的预处理和准备,我们使用了22,398条记录进行数据挖掘。识别患者的模型精度为0.717。由于较高的信息增益,年龄因素被放置在树的根节点中。ROC曲线表明了模型在识别患者和健康个体方面的作用。该曲线表明该模型具有较高的识别能力,特别是对健康人的识别能力更强。我们开发了一个使用决策树筛查T2 DM的模型,该模型不需要对T2 DM的诊断进行实验室测试。
The aim of this study was to examine a predictive model using features related to the diabetes type 2 risk factors. The data were obtained from a database in a diabetes control system in Tabriz, Iran. The data included all people referred for diabetes screening between 2009 and 2011. The features considered as “Inputs” were: age, sex, systolic and diastolic blood pressure, family history of diabetes, and body mass index (BMI). Moreover, we used diagnosis as “Class”. We applied the “Decision Tree” technique and “J48” algorithm in the WEKA (3.6.10 version) software to develop the model. After data preprocessing and preparation, we used 22,398 records for data mining. The model precision to identify patients was 0.717. The age factor was placed in the root node of the tree as a result of higher information gain. The ROC curve indicates the model function in identification of patients and those individuals who are healthy. The curve indicates high capability of the model, especially in identification of the healthy persons. We developed a model using the decision tree for screening T2DM which did not require laboratory tests for T2DM diagnosis.