Splitting strategies for post-selection inference

Splitting strategies for post-selection inference
复制标题

DOI:
10.1093/biomet/asac070
复制
发表时间:
2021-02
期刊:
影响因子:
2.7
通讯作者:
D. G. Rasines;G. A. Young
D. G. Rasines;G. A. Young
中科院分区:
数学2区
文献类型:
--
作者:
D. G. Rasines;G. A. Young

文献摘要

被引文献

相似文献

我们考虑在稀疏回归设置中为所选参数提供有效推断的问题。众所周知,由于选择步骤中产生的偏差,经典的回归工具在这种情况下可能是不可靠的。近年来,已经提出了许多方法来确保推论有效性。在这里,我们考虑了基于随机化响应向量的数据分裂的简单替代方法,该响应向量比前者更高的选择和推理能力,并且适用于任意选择规则。我们提供了这两种方法的理论和经验比较,并为随机方法提供了中心限制定理。我们的调查表明,权力的收益可能是可观的。
We consider the problem of providing valid inference for a selected parameter in a sparse regression setting. It is well known that classical regression tools can be unreliable in this context due to the bias generated in the selection step. Many approaches have been proposed in recent years to ensure inferential validity. Here, we consider a simple alternative to data splitting based on randomizing the response vector, which allows for higher selection and inferential power than the former and is applicable with an arbitrary selection rule. We provide a theoretical and empirical comparison of both methods and derive a central limit theorem for the randomization approach. Our investigations show that the gain in power can be substantial.