Prediction of Chronic Kidney Disease-A Machine Learning Perspective

Prediction of Chronic Kidney Disease-A Machine Learning Perspective
复制标题

DOI:
10.1109/access.2021.3053763
复制
发表时间:
2021-01-01
期刊:
影响因子:
3.9
通讯作者:
Bolshev, Vadim
Bolshev, Vadim
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chittora, Pankaj;Chaurasia, Sandeep;Bolshev, Vadim

文献摘要

被引文献

相似文献

慢性肾脏病是当今最严重的疾病之一,需要尽快进行正确的诊断。机器学习技术已经成为医疗的可靠手段。在机器学习分类算法的帮助下,医生可以及时发现疾病。从这个角度来看,慢性肾脏病的预测在本文中进行了讨论。慢性肾脏病数据集取自UCI存储库。本研究应用了人工神经网络、C5.0、卡方自动交互检测器、逻辑回归、惩罚L1和惩罚L2线性支持向量机、随机树等七种分类算法。重要特征选择技术也被应用到数据集。对于每个分类器,已经基于(i)全特征、(ii)基于相关性的特征选择、(iii)包装器方法特征选择、(iv)最小绝对收缩和选择算子回归、(v)具有最小绝对收缩和选择算子回归选择的特征的合成少数过采样技术、(vi)具有全特征的合成少数过采样技术计算了结果。结果表明,在全特征的合成少数过采样技术中,惩罚L2的LSVM的分类准确率最高,达到98.86%。沿着计算了准确率、精确率、召回率、F-度量、曲线下面积和GINI系数,并在图中显示了各种算法的比较结果。最小绝对收缩和选择算子回归选择的特征与合成少数过采样技术相比,全特征的合成少数过采样技术得到了最好的。在具有最小绝对收缩和选择算子选择特征的合成少数过采样技术中,再次线性支持向量机给出了最高的98.46%的准确率。沿着机器学习模型,一个深度神经网络已经应用于同一数据集,并且已经注意到深度神经网络达到了99.6%的最高准确度。
Chronic Kidney Disease is one of the most critical illness nowadays and proper diagnosis is required as soon as possible. Machine learning technique has become reliable for medical treatment. With the help of a machine learning classifier algorithms, the doctor can detect the disease on time. For this perspective, Chronic Kidney Disease prediction has been discussed in this article. Chronic Kidney Disease dataset has been taken from the UCI repository. Seven classifier algorithms have been applied in this research such as artificial neural network, C5.0, Chi-square Automatic interaction detector, logistic regression, linear support vector machine with penalty L1 & with penalty L2 and random tree. The important feature selection technique was also applied to the dataset. For each classifier, the results have been computed based on (i) full features, (ii) correlation-based feature selection, (iii) Wrapper method feature selection, (iv) Least absolute shrinkage and selection operator regression, (v) synthetic minority over-sampling technique with least absolute shrinkage and selection operator regression selected features, (vi) synthetic minority over-sampling technique with full features. From the results, it is marked that LSVM with penalty L2 is giving the highest accuracy of 98.86% in synthetic minority over-sampling technique with full features. Along with accuracy, precision, recall, F-measure, area under the curve and GINI coefficient have been computed and compared results of various algorithms have been shown in the graph. Least absolute shrinkage and selection operator regression selected features with synthetic minority over-sampling technique gave the best after synthetic minority over-sampling technique with full features. In the synthetic minority over-sampling technique with least absolute shrinkage and selection operator selected features, again linear support vector machine gave the highest accuracy of 98.46%. Along with machine learning models one deep neural network has been applied on the same dataset and it has been noted that deep neural network achieved the highest accuracy of 99.6%.