Evaluation of software for multiple imputation of semi-continuous data

Evaluation of software for multiple imputation of semi-continuous data
复制标题

DOI:
10.1177/0962280206074464
复制
发表时间:
2007-01-01
影响因子:
2.3
通讯作者:
Rivero-Arias, Oliver
Rivero-Arias, Oliver
中科院分区:
医学3区
文献类型:
--
作者:
Yu, L-M;Burton, Andrea;Rivero-Arias, Oliver

文献摘要

被引文献

相似文献

目前,人们普遍认为多重填补(MI)方法比单一填补方法能更好地处理缺失数据的不确定性。一些标准统计软件包,如SAS、R和STATA,都有执行MI的标准程序或用户编写的程序。对于大多数类型的数据,这些包的性能通常是可以接受的。然而,目前尚不清楚这些应用程序是否适合于输入具有较大比例零值的数据,从而导致半连续分布。此外,当需要保留数据的分布以供后续分析时,还不清楚使用这些应用程序是否合适。本文报告了一项模拟研究的结果,该研究旨在评估在这些统计包内处理半连续数据的MI过程的性能。使用来自大型随机临床试验的1060名参与者的完整资源使用数据作为模拟总体,从中获得500个引导样本并施加缺失数据。这项研究的结果显示,在输入半连续数据时,MI程序的表现存在差异。在决定哪个程序应该对这种类型的数据执行MI时应谨慎行事。
It is now widely accepted that multiple imputation (MI) methods properly handle the uncertainty of missing data over single imputation methods. Several standard statistical software packages, such as SAS, R and STATA, have standard procedures or user-written programs to perform MI. The performance of these packages is generally acceptable for most types of data. However, it is unclear whether these applications are appropriate for imputing data with a large proportion of zero values resulting in a semi-continuous distribution. In addition, it is not clear whether the use of these applications is suitable when the distribution of the data needs to be preserved for Subsequent analysis. This article reports the findings of a simulation study carried out to evaluate the performance of the MI procedures for handling semi-continuous data within these statistical packages. Complete resource use data on 1060 participants from a large randomized clinical trial were used as the simulation population from which 500 bootstrap samples were obtained and missing data imposed. The findings of this study showed differences in the performance of the MI programs when imputing semi-continuous data. Caution should be exercised when deciding which program should perform MI on this type of data.