Taking Parametric Assumptions Seriously: Arguments for the Use of Welch's F-test instead of the Classical F-test in One-Way ANOVA

Taking Parametric Assumptions Seriously: Arguments for the Use of Welch's F-test instead of the Classical F-test in One-Way ANOVA
复制标题

DOI:
10.5334/irsp.198
复制
发表时间:
2019-08-01
影响因子:
2.5
通讯作者:
Lakens, Daniel
Lakens, Daniel
中科院分区:
心理学3区
文献类型:
--
作者:
Delacre, Marie;Leys, Christophe;Lakens, Daniel

文献摘要

被引文献

相似文献

学生t检验和经典的F检验方差分析都依赖于这样的假设:两个或多个样本是独立的,独立且同分布的残差是正态的,组间具有相等的方差。我们把重点放在正态和方差相等的假设上,并认为这些假设在心理学领域往往是不现实的。通过对研究人员实践的分析,我们强调了目前对这些假设缺乏关注的情况。通过蒙特卡罗模拟,我们说明了在第一类错误率和统计功率不满足检验假设的情况下,对ANOVA执行经典的参数F检验的后果。在实际偏离等方差假设的情况下,经典的F检验可能会产生严重的偏差结果,并导致无效的统计推断。我们考察了F检验的两种常见替代方法,即W检验和Brown-Forsythe检验(F*检验)。我们的模拟表明,在一系列现实情景下,W-检验是一个更好的选择,因此我们建议在比较均值时默认使用W-检验。我们提供了一个详细的例子,解释了如何在SPSS和R中执行W检验。我们在实际建议中总结了我们的结论,研究人员可以使用这些建议来改进他们的统计实践。
Student's t-test and classical F-test ANOVA rely on the assumptions that two or more samples are independent, and that independent and identically distributed residuals are normal and have equal variances between groups. We focus on the assumptions of normality and equality of variances, and argue that these assumptions are often unrealistic in the field of psychology. We underline the current lack of attention to these assumptions through an analysis of researchers' practices. Through Monte Carlo simulations, we illustrate the consequences of performing the classic parametric F-test for ANOVA when the test assumptions are not met on the Type I error rate and statistical power. Under realistic deviations from the assumption of equal variances, the classic F-test can yield severely biased results and lead to invalid statistical inferences. We examine two common alternatives to the F-test, namely the Welch's ANOVA (W-test) and the Brown-Forsythe test (F*-test). Our simulations show that under a range of realistic scenarios, the W-test is a better alternative and we therefore recommend using the W-test by default when comparing means. We provide a detailed example explaining how to perform the W-test in SPSS and R. We summarize our conclusions in practical recommendations that researchers can use to improve their statistical practices.