Errors in Statistical Inference Under Model Misspecification: Evidence, Hypothesis Testing, and AIC

Errors in Statistical Inference Under Model Misspecification: Evidence, Hypothesis Testing, and AIC
复制标题

DOI:
10.3389/fevo.2019.00372
复制
发表时间:
2019-10-21
影响因子:
3
通讯作者:
Lele, Subhash R.
Lele, Subhash R.
中科院分区:
环境科学与生态学2区
文献类型:
--
作者:
Dennis, Brian;Ponciano, Jose Miguel;Lele, Subhash R.

文献摘要

被引文献

相似文献

即使在统计的频率论分支中,科学分析中进行统计推断的方法也已多样化,但比较却难以捉摸。我们对 Neyman-Pearson 假设检验、Fisher 显着性检验、信息标准和证据统计的性能进行分析和数值近似(Royall,1997)。最后一种方法以证据函数的形式实现:通过基于数据估计两个模型与生成过程(即真值)的相对距离来比较两个模型的统计数据(Lele,2004)。这个定义的一个结果是一个显着的特性,即误导性或弱证据的概率、类似于假设检验中的 1 类和 2 类错误的错误概率,随着样本量的增加,全部接近 0。我们对这些方法的比较主要集中在正确指定模型和错误指定模型时发生错误的频率,但也考虑了解释的难易程度。即使在模型错误指定的情况下,证据分析中的错误率也会随着样本量的增加而降低至 0。另一方面,内曼-皮尔逊测试在错误指定的情况下表现出很大的困难。实际 1 类和 2 类错误率可能小于、等于或大于名义率,具体取决于模型错误指定的性质。在某些合理的情况下,1 类错误的概率是样本量的增函数,甚至可以接近 1!相反,在模型错误指定的情况下,证据分析保留了理想的特性,即始终具有比较差模型更大的概率选择最佳模型,并且选择最佳模型的概率随着样本大小单调增加。我们表明,证据函数概念在统计和科学意义上都满足了生态学模型选择的表面目标,并且证据函数直观且易于掌握。我们发现一致的信息标准是证据函数,但 MSE 最小化(或有效)信息标准(例如 AIC、AICc、TIC)不是。 MSE 最小化标准的误差属性在证据函数的误差属性和 Neyman-Pearson 检验的误差属性之间切换,具体取决于所比较的模型。
The methods for making statistical inferences in scientific analysis have diversified even within the frequentist branch of statistics, but comparison has been elusive. We approximate analytically and numerically the performance of Neyman-Pearson hypothesis testing, Fisher significance testing, information criteria, and evidential statistics (Royall, 1997). This last approach is implemented in the form of evidence functions: statistics for comparing two models by estimating, based on data, their relative distance to the generating process (i.e., truth) (Lele, 2004). A consequence of this definition is the salient property that the probabilities of misleading or weak evidence, error probabilities analogous to Type 1 and Type 2 errors in hypothesis testing, all approach 0 as sample size increases. Our comparison of these approaches focuses primarily on the frequency with which errors are made, both when models are correctly specified, and when they are misspecified, but also considers ease of interpretation. The error rates in evidential analysis all decrease to 0 as sample size increases even under model misspecification. Neyman-Pearson testing on the other hand, exhibits great difficulties under misspecification. The real Type 1 and Type 2 error rates can be less, equal to, or greater than the nominal rates depending on the nature of model misspecification. Under some reasonable circumstances, the probability of Type 1 error is an increasing function of sample size that can even approach 1! In contrast, under model misspecification an evidential analysis retains the desirable properties of always having a greater probability of selecting the best model over an inferior one and of having the probability of selecting the best model increase monotonically with sample size. We show that the evidence function concept fulfills the seeming objectives of model selection in ecology, both in a statistical as well as scientific sense, and that evidence functions are intuitive and easily grasped. We find that consistent information criteria are evidence functions but the MSE minimizing (or efficient) information criteria (e.g., AIC, AICc, TIC) are not. The error properties of the MSE minimizing criteria switch between those of evidence functions and those of Neyman-Pearson tests depending on models being compared.