Investigating the impact of data normalization on classification performance

Investigating the impact of data normalization on classification performance
复制标题

DOI:
10.1016/j.asoc.2019.105524
复制
发表时间:
2020-12-01
影响因子:
8.7
通讯作者:
Singh, Birmohan
Singh, Birmohan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Singh, Dalwinder;Singh, Birmohan

文献摘要

被引文献

相似文献

数据归一化是预处理方法之一,其中数据被缩放或变换以使每个特征的贡献相等。机器学习算法的成功取决于获得分类问题的广义预测模型的数据质量。数据规范化对于提高数据质量和机器学习算法性能的重要性已经在许多研究中提出。但是,工作缺乏的特征选择和特征加权的方法,目前的研究趋势,机器学习,以提高性能。因此,本研究的目的是调查的影响,14个数据规范化方法的分类性能考虑全特征集,特征选择和特征加权。本文还提出了一种改进的蚁群算法,该算法利用最近邻分类器的参数搜索特征子集和最佳特征权值沿着。在21个公开的真实的和合成数据集上进行了实验,并根据准确率、特征减少百分比和运行时间对结果进行了分析。从结果中可以看出,没有一种方法优于其他方法。因此,我们提出了一套最好的和最坏的方法结合规范化程序和实证分析的结果。性能较好的是z-Score和Pareto Scaling用于完整的特征集和特征选择,以及tanh及其变体用于特征加权。表现最差的是均值中心、变量稳定性标度、中位数和中位数绝对偏差方法以及未归一化数据沿着。(C)2019由Elsevier B. V.出版
Data normalization is one of the pre-processing approaches where the data is either scaled or transformed to make an equal contribution of each feature. The success of machine learning algorithms depends upon the quality of the data to obtain a generalized predictive model of the classification problem. The importance of data normalization for improving data quality and subsequently the performance of machine learning algorithms has been presented in many studies. But, the work lacks for the feature selection and feature weighting approaches, a current research trend in machine learning for improving performance. Therefore, this study aims to investigate the impact of fourteen data normalization methods on classification performance considering full feature set, feature selection, and feature weighting. In this paper, we also present a modified Ant Lion optimization that search feature subsets and the best feature weights along with the parameter of Nearest Neighbor Classifier. Experiments are performed on 21 publicly available real and synthetic datasets, and results are analyzed based on the accuracy, the percentage of feature reduced and runtime. It has been observed from the results that no single method outperforms others. Therefore, we have suggested a set of the best and the worst methods combining the normalization procedure and empirical analysis of results. The better performers are z-Score and Pareto Scaling for the full feature set and feature selection, and tanh and its variant for feature weighting. The worst performers are Mean Centered, Variable Stability Scaling and Median and Median Absolute Deviation methods along with un-normalized data. (C) 2019 Published by Elsevier B.V.