Statistical Inference After Model Selection

Statistical Inference After Model Selection
复制标题

DOI:
10.1007/s10940-009-9077-7
复制
发表时间:
2010-06-01
影响因子:
3.6
通讯作者:
Zhao, Linda
Zhao, Linda
中科院分区:
法学1区
文献类型:
--
作者:
Berk, Richard;Brown, Lawrence;Zhao, Linda

文献摘要

被引文献

相似文献

传统的统计推断要求在分析数据之前知道数据是如何生成的模型。然而,在犯罪学和更广泛的社会科学中,通常会进行各种模型选择程序,然后进行统计检验和为“最终”模型计算置信区间。在本文中,我们研究这种做法,并显示他们是如何典型的误导。被估计的参数不再是很好的定义,和后模型选择抽样分布的混合物与属性是非常不同的,从传统上假设。置信区间和统计检验并没有发挥应有的作用。我们详细研究了负责的具体机制。我们还提供了一些建议,更好的做法,并显示通过刑事司法的例子,使用真实的数据如何正确的统计推断原则上可以获得。
Conventional statistical inference requires that a model of how the data were generated be known before the data are analyzed. Yet in criminology, and in the social sciences more broadly, a variety of model selection procedures are routinely undertaken followed by statistical tests and confidence intervals computed for a "final" model. In this paper, we examine such practices and show how they are typically misguided. The parameters being estimated are no longer well defined, and post-model-selection sampling distributions are mixtures with properties that are very different from what is conventionally assumed. Confidence intervals and statistical tests do not perform as they should. We examine in some detail the specific mechanisms responsible. We also offer some suggestions for better practice and show though a criminal justice example using real data how proper statistical inference in principle may be obtained.