The earth is flat (p > 0.05): significance thresholds and the crisis of unreplicable research.

The earth is flat (p > 0.05): significance thresholds and the crisis of unreplicable research.
复制标题

地球是平的(p > 0.05):重要性阈值和不可复制研究的危机。

DOI:
10.7717/peerj.3544
复制
发表时间:
2017
期刊:
影响因子:
2.7
通讯作者:
Roth T
Roth T
中科院分区:
生物学3区
文献类型:
--
作者:
Amrhein V;Korner-Nievergelt F;Roth T

文献摘要

参考文献

被引文献

相似文献

广泛使用“统计显著性”作为声称科学发现的许可证,导致科学过程的相当大的扭曲(根据美国统计协会)。我们回顾为什么降低p值到“显着”和“不显着”有助于使研究不可重现,或使他们似乎不可重现。一个主要的问题是,我们倾向于按面值取小p值,但不信任大p值的结果。在这两种情况下,p值几乎不能说明研究的可靠性,因为即使备择假设是正确的,它们也很难被复制。此外,显著性(p ≤ 0.05)几乎不可复制:在80%的良好统计功效下,两项研究将是“冲突”的,这意味着一项具有显著性,另一项不具有显著性,如果存在真实效应,则在三分之一的情况下。因此,复制不能仅仅因为不重要而被解释为失败。因此,许多明显的复制失败可能反映了基于重要性阈值的错误判断,而不是不可复制研究的危机。只有使用多项独立研究的累积证据,才能得出关于发现的可重复性和实际重要性的可靠结论。然而,应用显著性阈值使得累积知识不可靠。原因之一是,除了理想的统计功效,显著效应量将向上偏置。因此,解释夸大的重要结果而忽略不重要的结果将导致错误的结论。但是,目前追求重要性的动机导致了选择性报道和对非重要发现的出版偏见。数据挖掘,p-黑客和出版偏见应该通过删除固定的重要性阈值来解决。与已故罗纳德费舍尔的建议一致,p值应解释为对零假设的证据强度的分级测量。此外,较大的p值提供了一些反对零假设的证据,它们不能被解释为支持零假设,错误地得出“没有效果”的结论。必须从点估计值中获得与数据相容的可能真实效应量的信息,例如,样本平均值和区间估计值,例如置信区间。我们回顾了对较大p值解释的困惑如何可以追溯到现代统计学创始人之间的历史争议。我们进一步讨论了反对删除显著性阈值的潜在论点,例如,决策规则应该更严格,样本量可以减少,或者p值应该完全放弃。我们的结论是,无论我们使用什么统计推断方法,二分法阈值思维必须让位于非自动化的知情判断。
The widespread use of ‘statistical significance’ as a license for making a claim of a scientific finding leads to considerable distortion of the scientific process (according to the American Statistical Association). We review why degrading p-values into ‘significant’ and ‘nonsignificant’ contributes to making studies irreproducible, or to making them seem irreproducible. A major problem is that we tend to take small p-values at face value, but mistrust results with larger p-values. In either case, p-values tell little about reliability of research, because they are hardly replicable even if an alternative hypothesis is true. Also significance (p ≤ 0.05) is hardly replicable: at a good statistical power of 80%, two studies will be ‘conflicting’, meaning that one is significant and the other is not, in one third of the cases if there is a true effect. A replication can therefore not be interpreted as having failed only because it is nonsignificant. Many apparent replication failures may thus reflect faulty judgment based on significance thresholds rather than a crisis of unreplicable research. Reliable conclusions on replicability and practical importance of a finding can only be drawn using cumulative evidence from multiple independent studies. However, applying significance thresholds makes cumulative knowledge unreliable. One reason is that with anything but ideal statistical power, significant effect sizes will be biased upwards. Interpreting inflated significant results while ignoring nonsignificant results will thus lead to wrong conclusions. But current incentives to hunt for significance lead to selective reporting and to publication bias against nonsignificant findings. Data dredging, p-hacking, and publication bias should be addressed by removing fixed significance thresholds. Consistent with the recommendations of the late Ronald Fisher, p-values should be interpreted as graded measures of the strength of evidence against the null hypothesis. Also larger p-values offer some evidence against the null hypothesis, and they cannot be interpreted as supporting the null hypothesis, falsely concluding that ‘there is no effect’. Information on possible true effect sizes that are compatible with the data must be obtained from the point estimate, e.g., from a sample average, and from the interval estimate, such as a confidence interval. We review how confusion about interpretation of larger p-values can be traced back to historical disputes among the founders of modern statistics. We further discuss potential arguments against removing significance thresholds, for example that decision rules should rather be more stringent, that sample sizes could decrease, or that p-values should better be completely abandoned. We conclude that whatever method of statistical inference we use, dichotomous threshold thinking must give way to non-automated informed judgment.
DOI: 10.1177/0959354314525282
发表时间: 2014-04-01
影响因子: 1.2
作者:
Branch, Marc
通讯作者: Branch, Marc
DOI: 10.3389/fpsyg.2016.01247
发表时间: 2016-08-23
影响因子: 3.8
作者:
Badenes-Ribera, Laura;Frias-Navarro, Dolores;Longobardi, Claudio
通讯作者: Longobardi, Claudio
DOI: 10.1037/h0074554
发表时间: 1919-10-01
影响因子: 22.4
作者:
Boring, Edwin G.
通讯作者: Boring, Edwin G.
DOI: 10.1001/jama.2016.1952
发表时间: 2016-03-15
影响因子: 120.7
作者:
Chavalarias, David;Wallach, Joshua David;Ioannidis, John P. A.
通讯作者: Ioannidis, John P. A.
DOI: 10.2307/2982063
发表时间: 1980-01-01
影响因子: 2
作者:
BOX, GEP
通讯作者: BOX, GEP