Comparison of Performance of Data Imputation Methods for Numeric Dataset

Comparison of Performance of Data Imputation Methods for Numeric Dataset
复制标题

DOI:
10.1080/08839514.2019.1637138
复制
发表时间:
2019-07-06
影响因子:
2.8
通讯作者:
Ramanathan, Krishnan
Ramanathan, Krishnan
中科院分区:
计算机科学4区
文献类型:
--
作者:
Jadhav, Anil;Pramod, Dhanya;Ramanathan, Krishnan

文献摘要

被引文献

相似文献

数据缺失是研究人员和数据科学家面临的常见问题。因此,为了得到更好、更准确的数据分析结果,需要对它们进行适当的处理。本文的研究目的是为了更好地理解数据缺失机制、数据填入方法,并评估目前广泛使用的数字数据集数据填入方法的性能。它将帮助实践者和数据科学家在执行数据挖掘任务时,为数字数据集选择合适的数据输入方法。本文综合比较了均值拟合、中位数拟合、kNN拟合、预测均值拟合、贝叶斯线性回归(范数)、线性回归、非贝叶斯范数等7种数据拟合方法。Nob),随机抽样。我们使用了从UCI机器学习存储库中获得的五个不同的数字数据集来分析和比较数据输入方法的性能。采用归一化均方根误差(RMSE)方法对数据输入方法的性能进行了评估。分析结果表明,kNN插值方法优于其他方法。研究还发现,数据插入方法的性能与数据集和数据集中缺失值的百分比无关。
Missing data is common problem faced by researchers and data scientists. Therefore, it is required to handle them appropriately in order to get better and accurate results of data analysis. Objective of this research paper is to provide better understanding of data missingness mechanism, data imputation methods, and to assess performance of the widely used data imputation methods for numeric dataset. It will help practitioners and data scientists to select appropriate method of data imputation for numeric dataset while performing data mining task. In this paper, we comprehensively compare seven data imputation methods namely mean imputation, median imputation, kNN imputation, predictive mean matching, Bayesian Linear Regression (norm), Linear Regression, non-Bayesian (norm.nob), and random sample. We have used five different numeric datasets obtained from UCI machine learning repository for analyzing and comparing performance of the data imputation methods. Performance of the data imputation methods is assessed using Normalized Root Mean Square Error (RMSE) method. The results of analysis show that kNN imputation method outperforms the other methods. It has also been found that performance of the data imputation method is independent of the dataset and percentage of missing values in the dataset.