Beyond the Significance Test Ritual: What Is There?

Beyond the Significance Test Ritual: What Is There?
复制标题

除了重要性测试仪式之外:还有什么?

DOI:
10.1027/0044-3409.217.1.1
复制
发表时间:
2009
影响因子:
1.8
通讯作者:
P. Sedlmeier
P. Sedlmeier
中科院分区:
心理学3区
文献类型:
--
作者:
P. Sedlmeier

文献摘要

被引文献

相似文献

无头脑地使用零假设显著性检验--重要性检验的例行公事(例如Salsburg,1985)--长期以来一直受到批评。仪式的主要部分可以被描述为:一旦你收集了你的数据,试着驳斥你的零假设(例如,没有均值差异,零相关性,等等)。以自动化的方式。通常情况下,这一惯例还会得到“星级程序”的补充:如果P<.05,给你的结果分配一颗星(*),如果P<.01给你的结果分配两颗星(**),如果P<.001,你为自己赢得了三颗星(*)。如果你至少获得了一颗星,那么仪式就成功完成了;如果没有,你的结果就不值多少钱了。这些星星,或相应的数值,为著名的心理学期刊打开了大门,因此,这一仪式得到了强有力的加强。仪式没有坚实的理论基础;它似乎是作为罗纳德·A·费舍尔、杰西·内曼、埃贡·S·皮尔森和托马斯·贝耶斯(至少在仪式的一些变体中)的方法的一个令人费解的混合体而出现的(见Acree,1979;Gigerenzer&Murray,1987;Spielman,1974)。很长一段时间以来,关于它的有效性一直存在争议。然而,这场争论引起的争论并不局限于对上文概述的盲目程序的讨论,而是扩大到包括实验设计和抽样程序的问题,关于总体效应大小的假设(导致另一种假设的具体说明),在收集数据之前关于统计能力的审议,以及关于第一类和第二类错误的决定。已经有过几次这样的辩论,而且争议仍在继续(摘要见Balluerka,Gómez和Hidalgo,2005;Nickerson,2000;Sedlmeier,1999,附录C)。虽然有些人主张禁止显著性检验(如Hunter,1997),但作者通常认为,如果操作得当,显著性检验可能有一些价值(或至少不会造成伤害),但应该用其他更有信息量的数据分析方法来补充(或取而代之)(例如,Abelson,1995;Cohen,1994;Howard,Maxwell,&Fleming,2000;Loftus,1993;Nickerson,2000;Sedlmeier,1996;Wilkinson&Task Force on统计量推断,1999)。替代的数据分析技术在方法学家中已经广为人知几十年了,但这些知识主要收集在方法期刊上,到目前为止似乎对研究人员的实践几乎没有影响。我认为造成这种不令人满意的状况的主要原因有两个。首先,对于显著性检验结果的真正含义,似乎仍然存在相当大的误解(例如,Gordon,2001;Haller&Krauss,2002;Mittag&Thompson,2000;蒙特雷de-i-Bort,Pascual Llobell,&Frias-Navarro,2008)。第二,尽管在广为接受的摘要文章(如Wilkinson和工作队,1999年)中简要提到了备选方案,但很少以非技术性和详细的方式向非专业受众介绍这些备选方案。因此,原则上,研究人员可能愿意改变他们分析数据的方式,但学习替代方法所需的努力可能被认为太大了。本期特刊的主要目的是以非技术性的方式介绍这些替代数据分析方法的集合,该领域的专家对此进行了描述。在介绍特刊的内容之前,我将简要概述推理统计的理想状态,并讨论无意识意义检验和有心意义检验的区别。
The mindless use of null-hypothesis significance testing – the significance test ritual (e.g., Salsburg, 1985) – has long been criticized. The main component of the ritual can be characterized as follows: Once you have collected your data, try to refute your null hypothesis (e.g., no mean difference, zero correlation, etc.) in an automatized manner. Often the ritual is complemented by the “star procedure”: If p < .05, assign one star to your results (*), if p < .01 give two stars (**), and if p < .001 you have earned yourself three stars (***). If you have obtained at least one star, the ritual has been successfully performed; if not, your results are not worth much. The stars, or the corresponding numerical values, have been door-openers to prestigious psychology journals and, therefore, the ritual has received strong reinforcement. The ritual does not have a firm theoretical grounding; it seems to have arisen as a badly understood hybrid mixture of the approaches of Ronald A. Fisher, Jerzy Neyman, Egon S. Pearson, and (at least in some variations of the ritual) Thomas Bayes (see Acree, 1979; Gigerenzer & Murray, 1987; Spielman, 1974). For quite some time, there has been controversy over its usefulness. The debates arising from this controversy, however, have not been limited to discussions about the mindless procedure as sketched above, but have expanded to include the issues of experimental design and sampling procedures, assumptions about the size of population effects (leading to the specification of an alternative hypothesis), deliberations about statistical power before the data are collected, and decisions about Type I and Type II errors. There have been several such debates and the controversy is ongoing (for a summary see Balluerka, Gómez, & Hidalgo, 2005; Nickerson, 2000; Sedlmeier, 1999, Appendix C). Although there have been voices that argue for a ban on significance testing (e.g., Hunter, 1997), authors usually conclude that significance tests, if conducted properly, probably have some value (or at least do no harm) but should be complemented (or replaced) by other more informative ways of analyzing data (e.g., Abelson, 1995; Cohen, 1994; Howard, Maxwell, & Fleming, 2000; Loftus, 1993; Nickerson, 2000; Sedlmeier, 1996; Wilkinson & Task Force on Statistical Inference, 1999). Alternative data-analysis techniques have been wellknown among methodologists for decades but this knowledge, mainly collected in methods journals, seems to have had little impact on the practice of researchers to date. I see two main reasons for this unsatisfactory state of affairs. First, it appears that there is still a fair amount of misunderstanding about what the results of significance tests really mean (e.g., Gordon, 2001; Haller & Krauss, 2002; Mittag & Thompson, 2000; Monterde-i-Bort, Pascual Llobell, & Frias-Navarro, 2008). Second, although alternatives have been briefly mentioned in widely received summary articles (such as Wilkinson & Task Force on Statistical Inference, 1999), they have rarely been presented in a nontechnical and detailed manner to a nonspecialized audience. Thus, researchers might, in principle, be willing to change how they analyze data but the effort needed to learn about alternative methods might just be regarded as too great. The main aim of this special issue is to introduce a collection of these alternative data-analysis methods in a nontechnical way, described by experts in the field. Before introducing the contents of the special issue, I will briefly outline the ideal state of affairs in inference statistics and discuss the difference between mindless and mindful significance testing.