A cautionary tale on using imputation methods for inference in matched-pairs design

A cautionary tale on using imputation methods for inference in matched-pairs design
复制标题

DOI:
10.1093/bioinformatics/btaa082
复制
发表时间:
2020-05-15
期刊:
影响因子:
5.8
通讯作者:
Pauly, Markus
Pauly, Markus
中科院分区:
生物学3区
文献类型:
--
作者:
Ramosaj, Burim;Amro, Lubna;Pauly, Markus

文献摘要

被引文献

相似文献

动机:生物医学领域的归算程序已经变成了统计实践,因为进一步的分析可以忽略之前缺失值的存在。特别是,与更传统的MICE程序相比,随机森林等非参数imputation方案显示出良好的imputation性能。然而,它们对有效统计推断的影响迄今尚未得到分析。本文通过调查它们在推断不完全观察到的对的平均差异的有效性来缩小这一差距,同时反对它们与最近的一种方法,这种方法只适用于手头的给定观察。结果:我们的研究结果表明,(乘法)输入缺失值的机器学习方案可能会增加I型误差,或者在小到中等匹配对中导致相对较低的功率,即使在使用Rubin的多重输入规则修改测试统计后也是如此。除了广泛的模拟研究外,还考虑了来自乳腺癌基因研究的说明性数据示例。
Motivation: Imputation procedures in biomedical fields have turned into statistical practice, since further analyses can be conducted ignoring the former presence of missing values. In particular, non-parametric imputation schemes like the random forest have shown favorable imputation performance compared to the more traditionally used MICE procedure. However, their effect on valid statistical inference has not been analyzed so far. This article closes this gap by investigating their validity for inferring mean differences in incompletely observed pairs while opposing them to a recent approach that only works with the given observations at hand.Results: Our findings indicate that machine-learning schemes for (multiply) imputing missing values may inflate type I error or result in comparably low power in small-to-moderate matched pairs, even after modifying the test statistics using Rubin's multiple imputation rule. In addition to an extensive simulation study, an illustrative data example from a breast cancer gene study has been considered.