Analyzing data sets with missing data: An empirical evaluation of imputation methods and likelihood-based methods

Analyzing data sets with missing data: An empirical evaluation of imputation methods and likelihood-based methods
复制标题

DOI:
10.1109/32.965340
复制
发表时间:
2001-11-01
影响因子:
7.4
通讯作者:
Olsson, UH
Olsson, UH
中科院分区:
计算机科学1区
文献类型:
--
作者:
Myrtveit, I;Stensrud, E;Olsson, UH

文献摘要

被引文献

相似文献

在构建努力预测模型的数据集中,经常会遇到数据缺失的问题。到目前为止,通常的做法是忽略缺少数据的观测结果。这可能导致有偏差的预测模型。在本文中,我们评估了软件成本建模背景下的四种缺失数据技术(MDTs):列表删除(LID),平均imputation (MI),相似响应模式imputation (SRPI)和完全信息最大似然(FIML)。我们将mdt应用于ERP数据集,然后使用结果数据集构建基于回归的预测模型。评价表明,只有当数据不是完全随机丢失(MCAR)时,才适用于FIML。与FIML不同,除非数据是MCAR,否则在LD、MI和SRPI数据集上构建的预测模型将存在偏差。此外,与LID相比,MI和SRPI似乎只有在结果LID数据集太小而无法构建有意义的基于回归的预测模型时才合适。
Missing data are often encountered in data sets used to construct effort prediction models. Thus far, the common practice has been to ignore observations with missing data. This may result in biased prediction models. In this paper, we evaluate four missing data techniques (MDTs) in the context of software cost modeling: listwise deletion (LID), mean imputation (MI), similar response pattern imputation (SRPI), and full information maximum likelihood (FIML). We apply the MDTs to an ERP data set, and thereafter construct regression-based prediction models using the resulting data sets. The evaluation suggests that only FIML is appropriate when the data are not missing completely at random (MCAR). Unlike FIML, prediction models constructed on LD, MI and SRPI data sets will be biased unless the data are MCAR. Furthermore, compared to LID, MI and SRPI seem appropriate only if the resulting LID data set is too small to enable the construction of a meaningful regression-based prediction model.