Predicting disease risks from highly imbalanced data using random forest

Predicting disease risks from highly imbalanced data using random forest
复制标题

DOI:
10.1186/1472-6947-11-51
复制
发表时间:
2011-07-29
影响因子:
3.5
通讯作者:
Popescu, Mihail
Popescu, Mihail
中科院分区:
医学3区
文献类型:
--
作者:
Khalilia, Mohammed;Chakraborty, Sounak;Popescu, Mihail

文献摘要

被引文献

相似文献

工作背景:我们提出了一种方法,利用医疗成本和利用项目(HCUP)数据集预测疾病风险的个人的基础上,他们的医疗诊断史。所提出的方法可以被纳入各种应用程序,如风险管理,量身定制的健康沟通和决策支持系统在healthcare.Methods:我们采用了国家住院病人样本(NIS)的数据,这是公开的,通过医疗保健成本和利用项目(HCUP),训练随机森林分类疾病预测。由于HCUP数据是高度不平衡的,我们采用了集成学习方法的基础上重复随机子采样。该技术将训练数据划分为多个子样本,同时确保每个子样本完全平衡。我们比较了支持向量机(SVM),bagging,boosting和RF的性能来预测八种慢性疾病的风险。总体而言,RF集成学习方法在受试者工作特征(ROC)曲线(AUC)下的面积方面优于SVM、bagging和boosting。此外,RF具有计算每个变量的重要性在classificationprocess.Conclusions的优势:在重复随机子采样与RF相结合,我们能够克服类不平衡的问题,并取得了可喜的成果。使用国家HCUP数据集,我们预测了8种疾病类别,平均AUC为88.79%。
Background: We present a method utilizing Healthcare Cost and Utilization Project (HCUP) dataset for predicting disease risk of individuals based on their medical diagnosis history. The presented methodology may be incorporated in a variety of applications such as risk management, tailored health communication and decision support systems in healthcare.Methods: We employed the National Inpatient Sample (NIS) data, which is publicly available through Healthcare Cost and Utilization Project (HCUP), to train random forest classifiers for disease prediction. Since the HCUP data is highly imbalanced, we employed an ensemble learning approach based on repeated random sub-sampling. This technique divides the training data into multiple sub-samples, while ensuring that each sub-sample is fully balanced. We compared the performance of support vector machine (SVM), bagging, boosting and RF to predict the risk of eight chronic diseases.Results: We predicted eight disease categories. Overall, the RF ensemble learning method outperformed SVM, bagging and boosting in terms of the area under the receiver operating characteristic (ROC) curve (AUC). In addition, RF has the advantage of computing the importance of each variable in the classification process.Conclusions: In combining repeated random sub-sampling with RF, we were able to overcome the class imbalance problem and achieve promising results. Using the national HCUP data set, we predicted eight disease categories with an average AUC of 88.79%.