Statistical Evidence in Experimental Psychology: An Empirical Comparison Using 855 t Tests

Statistical Evidence in Experimental Psychology: An Empirical Comparison Using 855 t Tests
复制标题

DOI:
10.1177/1745691611406923
复制
发表时间:
2011-05-01
影响因子:
12.6
通讯作者:
Wagenmakers, Eric-Jan
Wagenmakers, Eric-Jan
中科院分区:
心理学1区
文献类型:
--
作者:
Wetzels, Ruud;Matzke, Dora;Wagenmakers, Eric-Jan

文献摘要

被引文献

相似文献

心理学中的统计推断传统上严重依赖于p值显著性检验。然而,这种从数据中得出结论的方法受到了广泛的批评,并提出了两种补救措施。第一个建议是用证据的补充度量来补充p值,比如效应大小。第二种是用贝叶斯证据度量(如贝叶斯因子)代替推理。作者使用855个最近发表的心理学t检验,对p值、效应大小和默认贝叶斯因子作为统计证据的度量进行了实际比较。这种比较产生了两个主要结果。首先,尽管p值和默认贝叶斯因子几乎总是在数据更好地支持哪个假设的问题上达成一致,但这些指标往往在这种支持的强度上存在分歧;对于70%的p值落在。01和。05,默认的贝叶斯因子表明证据只是轶事。其次,效应大小可以为p值和默认贝叶斯因子提供额外的证据。作者得出结论,贝叶斯方法是相对谨慎的,防止研究人员高估证据,以支持一种效果。
Statistical inference in psychology has traditionally relied heavily on p-value significance testing. This approach to drawing conclusions from data, however, has been widely criticized, and two types of remedies have been advocated. The first proposal is to supplement p values with complementary measures of evidence, such as effect sizes. The second is to replace inference with Bayesian measures of evidence, such as the Bayes factor. The authors provide a practical comparison of p values, effect sizes, and default Bayes factors as measures of statistical evidence, using 855 recently published t tests in psychology. The comparison yields two main results. First, although p values and default Bayes factors almost always agree about what hypothesis is better supported by the data, the measures often disagree about the strength of this support; for 70% of the data sets for which the p value falls between .01 and .05, the default Bayes factor indicates that the evidence is only anecdotal. Second, effect sizes can provide additional evidence to p values and default Bayes factors. The authors conclude that the Bayesian approach is comparatively prudent, preventing researchers from overestimating the evidence in favor of an effect.