Effect of Data Scaling Methods on Machine Learning Algorithms and Model Performance

Effect of Data Scaling Methods on Machine Learning Algorithms and Model Performance
复制标题

DOI:
10.3390/technologies9030052
复制
发表时间:
2021-09-01
期刊:
影响因子:
3.6
通讯作者:
Siddique, Zahed
Siddique, Zahed
中科院分区:
其他
文献类型:
--
作者:
Ahsan, Md Manjurul;Mahmud, M. A. Parvez;Siddique, Zahed

文献摘要

被引文献

相似文献

心脏病是世界各地高死亡率的主要原因之一,需要复杂而昂贵的诊断过程。近年来,许多文献已经证明机器学习方法是有效诊断心脏病患者的一个机会。然而,与数据集相关的挑战,例如缺失数据、不一致的数据和混合数据(包含不一致的数值和分类缺失数据)通常是医学诊断中的障碍。这种不一致性导致了更高的错误预测和误导结果的可能性。数据预处理步骤,如特征约简,数据转换和数据缩放,以形成一个标准的样本集,这些措施在减少最终预测的不准确性发挥了至关重要的作用。本文旨在评估11种机器学习(ML)算法-逻辑回归(LR),线性判别分析(LDA),K-最近邻(KNN),分类和回归树(CART),朴素贝叶斯(NB),支持向量机(SVM),XGBoost(XGB),随机森林分类器(RF),梯度提升(GB),AdaBoost(AB),额外的树分类器(ET)-和六种不同的数据缩放方法-归一化(NR),Standscale(SS),MinMax(MM),MaxAbs(MA),Robust Scaler(RS)和Quantile Transformer(QT),数据集包括心脏病患者的信息。结果表明,CART,沿着与RS或QT,优于所有其他ML算法100%的准确率,100%的精度,99%的召回率,和100%F1得分。研究结果表明,该模型的性能取决于数据缩放方法。
Heart disease, one of the main reasons behind the high mortality rate around the world, requires a sophisticated and expensive diagnosis process. In the recent past, much literature has demonstrated machine learning approaches as an opportunity to efficiently diagnose heart disease patients. However, challenges associated with datasets such as missing data, inconsistent data, and mixed data (containing inconsistent missing data both as numerical and categorical) are often obstacles in medical diagnosis. This inconsistency led to a higher probability of misprediction and a misled result. Data preprocessing steps like feature reduction, data conversion, and data scaling are employed to form a standard dataset-such measures play a crucial role in reducing inaccuracy in final prediction. This paper aims to evaluate eleven machine learning (ML) algorithms-Logistic Regression (LR), Linear Discriminant Analysis (LDA), K-Nearest Neighbors (KNN), Classification and Regression Trees (CART), Naive Bayes (NB), Support Vector Machine (SVM), XGBoost (XGB), Random Forest Classifier (RF), Gradient Boost (GB), AdaBoost (AB), Extra Tree Classifier (ET)-and six different data scaling methods-Normalization (NR), Standscale (SS), MinMax (MM), MaxAbs (MA), Robust Scaler (RS), and Quantile Transformer (QT) on a dataset comprising of information of patients with heart disease. The result shows that CART, along with RS or QT, outperforms all other ML algorithms with 100% accuracy, 100% precision, 99% recall, and 100% F1 score. The study outcomes demonstrate that the model's performance varies depending on the data scaling method.