Practical advice on how to impute continuous data when the ultimate interest centers on dichotomized outcomes through pre-specified thresholds

Practical advice on how to impute continuous data when the ultimate interest centers on dichotomized outcomes through pre-specified thresholds
复制标题

DOI:
10.1080/03610910701418424
复制
发表时间:
2007-01-01
影响因子:
0.9
通讯作者:
Demirtas, Hakan
Demirtas, Hakan
中科院分区:
数学4区
文献类型:
--
作者:
Demirtas, Hakan

文献摘要

被引文献

相似文献

在过去的二十年中,多元正态假设下的多重插补通常被认为是处理不完全连续数据的一种可行的基于模型的方法。在应用研究中,特别是在医学和社会科学中,以连续的规模进行测量,并通过特定学科的阈值对二分法版本产生最终兴趣的情况并不少见。在实践中,研究人员通常倾向于在高斯插补模型下对连续结果进行缺失值插补,然后通过普遍接受的截止点对它们进行二分。另一种策略是在使用饱和多项式结构的对数线性插补模型下进行二分后创建多重插补数据集。在这项工作中,这两种插补方法的性能进行了检查,在相当广泛的模拟不完整的数据集,表现出不同的分布特征,如偏度和多峰。探讨了效率和准确性措施的行为,以确定程序正常工作的程度。得出的结论是,进行对数线性插补之前的二分法应该是首选的方法,除了少数特殊情况。我建议研究人员使用非典型的第二种策略,只要兴趣集中在通过基础连续测量获得的二进制量上。一个可能的解释是,高斯模型不适应的不稳定/特殊方面可能在这个特定的缺失数据设置中转化为表现更好的离散趋势。这个前提胜过了连续变量固有地携带更多信息的断言,导致了一个反直觉的,但对从业者来说可能有用的结果。
Multiple imputation under the multivariate normality assumption has often been regarded as a viable model-based approach in dealing with incomplete continuous data in the last two decades. A situation where the measurements are taken on a continuous scale with an ultimate interest in dichotomized versions through discipline-specific thresholds is not uncommon in applied research, especially in medical and social sciences. In practice, researchers generally tend to impute missing values for continuous outcomes under a Gaussian imputation model, and then dichotomize them via commonly-accepted cut-off points. An alternative strategy is creating multiply imputed data sets after dichotomization under a log-linear imputation model that uses a saturated multinomial structure. In this work, the performances of the two imputation methods were examined on a fairly wide range of simulated incomplete data sets that exhibit varying distributional characteristics such as skewness and multimodality. Behavior of efficiency and accuracy measures was explored to determine the extent to which the procedures work properly. The conclusion drawn is that dichotomization before carrying out a log-linear imputation should be the preferred approach except for a few special cases. I recommend that researchers use the atypical second strategy whenever the interest centers on binary quantities that are obtained through underlying continuous measurements. A possible explanation is that erratic/idiosyncratic aspects that are not accommodated by a Gaussian model are probably transformed into better-behaving discrete trends in this particular missing-data setting. This premise outweighs the assertion that continuous variables inherently carry more information, leading to a counter-intuitive, but potentially useful result for practitioners.