Using multiple imputation to estimate missing data in meta-regression

Using multiple imputation to estimate missing data in meta-regression
复制标题

DOI:
10.1111/2041-210x.12322
复制
发表时间:
2015-02-01
影响因子:
6.6
通讯作者:
Murray, Dennis L.
Murray, Dennis L.
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Ellington, E. Hance;Bastille-Rousseau, Guillaume;Murray, Dennis L.

文献摘要

被引文献

相似文献

在生态学和进化论方面越来越需要科学的综合。在许多情况下,元分析技术可以用来补充这种综合。然而,缺失数据对于任何综合努力都是一个严重的问题,并且可能损害这些和其他学科的荟萃分析的完整性。目前,在生态学的元分析数据集和不同的补救措施,这一问题的有效性缺失数据的患病率还没有得到充分的量化。我们基于对实验和观测数据的文献综述生成了元分析数据集,发现缺失数据在元分析生态数据集中很普遍。然后,我们测试了完整病例删除(数据缺失时广泛使用的方法)和多重插补(数据恢复的替代方法)的性能,并使用已发布的荟萃回归数据集在各种模拟条件下评估模型偏倚,精度和多模型排名。我们发现,完全删除的情况下,导致有偏见和不精确的系数估计,并产生指定的模型。相比之下,多重插补提供了无偏的参数估计值,精度损失很小。然而,多重插补的性能取决于缺失数据的类型。当缺失值是加权变量时,它表现得最好,但当缺失值是预测变量时,性能好坏参半。当插补原始数据时,多重插补表现不佳,然后将其用于计算效应量和加权变量。我们的结论是,完全的情况下删除不应该使用元回归和多重插补有可能成为一个不可或缺的工具,在生态学和进化的元回归。但是,我们建议用户在实施多重插补以恢复实际缺失数据之前,通过在其数据的子集上模拟缺失数据来评估多重插补的性能。
There is a growing need for scientific synthesis in ecology and evolution. In many cases, meta-analytic techniques can be used to complement such synthesis. However, missing data are a serious problem for any synthetic efforts and can compromise the integrity of meta-analyses in these and other disciplines. Currently, the prevalence of missing data in meta-analytic data sets in ecology and the efficacy of different remedies for this problem have not been adequately quantified. We generated meta-analytic data sets based on literature reviews of experimental and observational data and found that missing data were prevalent in meta-analytic ecological data sets. We then tested the performance of complete case removal (a widely used method when data are missing) and multiple imputation (an alternative method for data recovery) and assessed model bias, precision and multimodel rankings under a variety of simulated conditions using published meta-regression data sets. We found that complete case removal led to biased and imprecise coefficient estimates and yielded poorly specified models. In contrast, multiple imputation provided unbiased parameter estimates with only a small loss in precision. The performance of multiple imputation, however, was dependent on the type of data missing. It performed best when missing values were weighting variables, but performance was mixed when missing values were predictor variables. Multiple imputation performed poorly when imputing raw data which were then used to calculate effect size and the weighting variable. We conclude that complete case removal should not be used in meta-regression and that multiple imputation has the potential to be an indispensable tool for meta-regression in ecology and evolution. However, we recommend that users assess the performance of multiple imputation by simulating missing data on a subset of their data before implementing it to recover actual missing data.