A Bayesian Perspective on the Reproducibility Project: Psychology.

A Bayesian Perspective on the Reproducibility Project: Psychology.
复制标题

DOI:
10.1371/journal.pone.0149794
复制
发表时间:
2016
期刊:
影响因子:
3.7
通讯作者:
Vandekerckhove J
Vandekerckhove J
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Etz A;Vandekerckhove J

文献摘要

被引文献

相似文献

我们重新审视了开放科学合作组织近期的“可重复性项目:心理学”的结果。我们计算了贝叶斯因子——一个可用于表达对一个假设以及零假设的比较证据的量——针对原始论文的一个大子集(N = 72)及其相应的重复实验。在我们的计算中,我们考虑到发表偏倚可能扭曲了最初发表的结果这一情况。总体而言,就所提供的证据量而言,75%的研究给出了性质相似的结果。然而,证据往往很薄弱(即贝叶斯因子<10)。大多数研究(64%)在原始研究或重复实验中,对于零假设或备择假设都没有提供强有力的证据,并且没有重复实验提供支持零假设的有力证据。在所有原始论文提供了有力证据但重复实验没有的情况(15%)中,重复实验的样本量比原始研究小。在重复实验提供了有力证据但原始研究没有的情况(10%)中,重复实验的样本量更大。我们得出结论,可重复性项目在重复许多目标效应方面明显失败,可以通过心理文献中由于样本量小和发表偏倚导致的效应量高估(或对零假设的反证高估)来充分解释。我们进一步得出结论,传统的样本量是不够的,更广泛地采用贝叶斯方法是可取的。
We revisit the results of the recent Reproducibility Project: Psychology by the Open Science Collaboration. We compute Bayes factors—a quantity that can be used to express comparative evidence for an hypothesis but also for the null hypothesis—for a large subset (N = 72) of the original papers and their corresponding replication attempts. In our computation, we take into account the likely scenario that publication bias had distorted the originally published results. Overall, 75% of studies gave qualitatively similar results in terms of the amount of evidence provided. However, the evidence was often weak (i.e., Bayes factor < 10). The majority of the studies (64%) did not provide strong evidence for either the null or the alternative hypothesis in either the original or the replication, and no replication attempts provided strong evidence in favor of the null. In all cases where the original paper provided strong evidence but the replication did not (15%), the sample size in the replication was smaller than the original. Where the replication provided strong evidence but the original did not (10%), the replication sample size was larger. We conclude that the apparent failure of the Reproducibility Project to replicate many target effects can be adequately explained by overestimation of effect sizes (or overestimation of evidence against the null hypothesis) due to small sample sizes and publication bias in the psychological literature. We further conclude that traditional sample sizes are insufficient and that a more widespread adoption of Bayesian methods is desirable.