Dealing with missing data in a multi-question depression scale: a comparison of imputation methods.

Dealing with missing data in a multi-question depression scale: a comparison of imputation methods.
复制标题

DOI:
10.1186/1471-2288-6-57
复制
发表时间:
2006-12-13
影响因子:
4
通讯作者:
Ghali WA
Ghali WA
中科院分区:
医学3区
文献类型:
--
作者:
Shrive FM;Stuart H;Quan H;Ghali WA

文献摘要

被引文献

相似文献

数据缺失给许多研究项目带来了挑战。这个问题在使用自我报告量表的研究中往往很明显,而关于在这种情况下处理缺失数据的不同策略的文献很少。本研究的目的是比较六种不同的方法处理Zung抑郁自评量表中缺失数据的方法。来自手术结果研究的1580名参与者完成了抑郁自评量表。SDS是一个20个问题的量表,受访者通过在每个问题上圈出1到4的值来完成。这些回答的总和被计算出来,当受访者的总分超过40分时,他们被归类为表现出抑郁症状。通过随机选择问题来模拟缺失的值,这些问题的值随后被删除(完全随机模拟的缺失)。此外,还完成了随机缺失和非随机缺失的仿真。然后考虑了六种归因法:1)多重归因法,2)一次回归法,3)个人平均数,4)总体平均数,5)参与者先前的反应,以及6)随机选择1到4之间的值。对于每种方法,将所计算的平均抑郁自评量表得分和标准差与总体统计数据进行比较。计算了Spearman相关系数、错别率和Kappa统计量。当10%的值丢失时,除随机选择外,所有的推算方法都产生大于0.80的Kappa统计量,表示接近完美的一致性。MI以高Kappa统计量(0.89)产生最有效的归因值,尽管单一回归和个体平均归因值也产生了良好的结果。当丢失信息的百分比增加到30%时,或者当引入不平衡的丢失数据时,MI保持较高的Kappa统计量。个体平均值和一次回归方法得出的Kappas在“基本一致”的范围内(分别为0.76和0.74)。在我们为SDS评估的大多数缺失数据场景中,多重填补是处理缺失数据的最准确方法。推算个人的平均值也是处理缺失数据的一种适当和简单的方法,可能对大多数医学读者来说更容易理解。当面临缺失数据时,研究人员应该考虑进行像这样的方法评估。最优的方法应该平衡有效性、读者的易解性和研究团队的分析专业知识。
Missing data present a challenge to many research projects. The problem is often pronounced in studies utilizing self-report scales, and literature addressing different strategies for dealing with missing data in such circumstances is scarce. The objective of this study was to compare six different imputation techniques for dealing with missing data in the Zung Self-reported Depression scale (SDS). 1580 participants from a surgical outcomes study completed the SDS. The SDS is a 20 question scale that respondents complete by circling a value of 1 to 4 for each question. The sum of the responses is calculated and respondents are classified as exhibiting depressive symptoms when their total score is over 40. Missing values were simulated by randomly selecting questions whose values were then deleted (a missing completely at random simulation). Additionally, a missing at random and missing not at random simulation were completed. Six imputation methods were then considered; 1) multiple imputation, 2) single regression, 3) individual mean, 4) overall mean, 5) participant's preceding response, and 6) random selection of a value from 1 to 4. For each method, the imputed mean SDS score and standard deviation were compared to the population statistics. The Spearman correlation coefficient, percent misclassified and the Kappa statistic were also calculated. When 10% of values are missing, all the imputation methods except random selection produce Kappa statistics greater than 0.80 indicating 'near perfect' agreement. MI produces the most valid imputed values with a high Kappa statistic (0.89), although both single regression and individual mean imputation also produced favorable results. As the percent of missing information increased to 30%, or when unbalanced missing data were introduced, MI maintained a high Kappa statistic. The individual mean and single regression method produced Kappas in the 'substantial agreement' range (0.76 and 0.74 respectively). Multiple imputation is the most accurate method for dealing with missing data in most of the missind data scenarios we assessed for the SDS. Imputing the individual's mean is also an appropriate and simple method for dealing with missing data that may be more interpretable to the majority of medical readers. Researchers should consider conducting methodological assessments such as this one when confronted with missing data. The optimal method should balance validity, ease of interpretability for readers, and analysis expertise of the research team.