Predicting Chronic Kidney Disease Using Hybrid Machine Learning Based on Apache Spark.

Predicting Chronic Kidney Disease Using Hybrid Machine Learning Based on Apache Spark.
复制标题

DOI:
10.1155/2022/9898831
复制
发表时间:
2022
影响因子:
--
通讯作者:
Goher N
Goher N
中科院分区:
工程技术3区
文献类型:
--
作者:
Abdel-Fattah MA;Othman NA;Goher N

文献摘要

参考文献

被引文献

相似文献

慢性肾脏病(CKD)已成为一种广泛流行的疾病。它与各种严重风险有关,如心血管疾病、高风险和终末期肾病,这些风险可以通过早期发现和治疗有这种疾病危险的人来避免。机器学习算法是医学科学家在疾病开始阶段准确诊断疾病的重要帮助来源。最近,大数据平台与机器学习算法相结合,为医疗保健增加了价值。因此,本文提出了混合机器学习技术,包括基于大数据平台(Apache Spark)的特征选择方法和机器学习分类算法,用于检测慢性肾脏疾病(CKD)。特征选择技术,即Relief-F和卡方特征选择方法,被用来选择重要的功能。在这项研究中使用了六种机器学习分类算法:决策树(DT),逻辑回归(LR),朴素贝叶斯(NB),随机森林(RF),支持向量机(SVM)和顺应性提升树(GBT分类器)作为集成学习算法。四种评价方法,即准确率,精确度,召回率,和F1-措施,应用于验证结果。对于每种算法,交叉验证的结果和测试结果已经计算的基础上的全功能,功能选择的Relief-F,卡方特征选择方法选择的功能。结果表明,SVM,DT和GBT分类器与选定的功能取得了最好的性能在100%的准确率。总体而言,Relief-F的选择特征优于完整特征和卡方选择特征。
Chronic kidney disease (CKD) has become a widespread disease among people. It is related to various serious risks like cardiovascular disease, heightened risk, and end-stage renal disease, which can be feasibly avoidable by early detection and treatment of people in danger of this disease. The machine learning algorithm is a source of significant assistance for medical scientists to diagnose the disease accurately in its outset stage. Recently, Big Data platforms are integrated with machine learning algorithms to add value to healthcare. Therefore, this paper proposes hybrid machine learning techniques that include feature selection methods and machine learning classification algorithms based on big data platforms (Apache Spark) that were used to detect chronic kidney disease (CKD). The feature selection techniques, namely, Relief-F and chi-squared feature selection method, were applied to select the important features. Six machine learning classification algorithms were used in this research: decision tree (DT), logistic regression (LR), Naive Bayes (NB), Random Forest (RF), support vector machine (SVM), and Gradient-Boosted Trees (GBT Classifier) as ensemble learning algorithms. Four methods of evaluation, namely, accuracy, precision, recall, and F1-measure, were applied to validate the results. For each algorithm, the results of cross-validation and the testing results have been computed based on full features, the features selected by Relief-F, and the features selected by chi-squared feature selection method. The results showed that SVM, DT, and GBT Classifiers with the selected features had achieved the best performance at 100% accuracy. Overall, Relief-F's selected features are better than full features and the features selected by chi-square.
DOI: 10.1109/access.2021.3053763
发表时间: 2021-01-01
期刊: IEEE ACCESS
影响因子: 3.9
作者:
Chittora, Pankaj;Chaurasia, Sandeep;Bolshev, Vadim
通讯作者: Bolshev, Vadim
DOI: 10.1159/000489897
发表时间: 2018-01-01
期刊: NEPHRON
影响因子: 2.5
作者:
Bikbov, Boris;Perico, Norberto;Remuzzi, Giuseppe
通讯作者: Remuzzi, Giuseppe
DOI: 10.1007/s10115-017-1145-y
发表时间: 2018-10-01
影响因子: 2.7
作者:
Palma-Mendoza, Raul-Jose;Rodriguez, Daniel;de-Marcos, Luis
通讯作者: de-Marcos, Luis
DOI: 10.1109/tasl.2013.2244083
发表时间: 2013-05-01
影响因子: --
作者:
Deng, Li;Li, Xiao
通讯作者: Li, Xiao
DOI: 10.1016/j.tust.2020.103677
发表时间: 2021-02-01
影响因子: 6.9
作者:
Huang, M. Q.;Ninic, J.;Zhang, Q. B.
通讯作者: Zhang, Q. B.