Estimating missing data: an iterative regression approach

Estimating missing data: an iterative regression approach
复制标题

DOI:
10.1006/jhev.2000.0418
复制
发表时间:
2000-09-01
影响因子:
3.2
通讯作者:
Benfer, RA
Benfer, RA
中科院分区:
地球科学2区
文献类型:
--
作者:
Holt, B;Benfer, RA

文献摘要

被引文献

相似文献

数据缺失的问题在所有科学领域都很常见。存在各种估计数据集中缺失值的方法,例如删除病例、插入样本均值和线性回归。每种方法都存在方法本身或缺失数据模式的性质所固有的问题。我们报告了一种方法,(1)在应用中更一般,(2)提供比传统方法更好的估计,例如一步回归模型是通用的,因为它可以应用于奇异矩阵,例如小数据集或包含虚拟或索引变量的数据集。该模型的优势在于它使用自助法迭代地构建回归方程。变量的回归估计值的精度随着预测变量的回归估计值的提高而提高。我们说明了这种方法与欧洲旧石器时代晚期和中石器时代人类颅后遗骸的测量,以及一组灵长类动物的人体测量数据。首先,使用灵长类动物数据集的模拟测试涉及随机将20%的值变为“缺失”。在每种情况下,第一次迭代产生的估计值明显优于其他估计技术。其次,我们将我们的方法应用于人类颅后测量的不完整集。MISDAT估计总是比用均值替换缺失数据和经典的多元回归更好。与经典的多元回归一样,MISDAT在平方多个相关值接近待估计测量的可靠性时执行,例如,高于约0.8。(C)北京大学出版社.
The problem of missing data is common in all fields of science. Various methods of estimating missing values in a dataset exist, such as deletion of cases, insertion of sample mean, and linear regression. Each approach presents problems inherent in the method itself or in the nature of the pattern of missing data. We report a method that (1) is more general in application and (2) provides better estimates than traditional approaches, such as one-step regression The model is general in that it may be applied to singular matrices, such as small datasets or those that contain dummy or index variables. The strength of the model is that it builds a regression equation iteratively, using a bootstrap method. The precision of the regressed estimates of a variable increases as regressed estimates of the predictor variables improve. We illustrate this method with a set of measurements of European Upper Paleolithic and Mesolithic human postcranial remains, as well as a set of primate anthropometric data. First, simulation tests using the primate data set involved randomly turning 20% of the values to "missing". In each case, the first iteration produced significantly better estimates than other estimating techniques. Second, we applied our method to the incomplete set of human postcranial measurements. MISDAT estimates always perform better than replacement of missing data by means and better than classical multiple regression. As with classical multiple regression, MISDAT performs when squared multiple correlation values approach the reliability of the measurement to be estimated, e.g., above about 0.8. (C) 2000 Academic Press.