Many Labs 5: Testing Pre-Data-Collection Peer Review as an Intervention to Increase Replicability

Many Labs 5: Testing Pre-Data-Collection Peer Review as an Intervention to Increase Replicability
复制标题

DOI:
10.1177/2515245920958687
复制
发表时间:
2020-09-01
影响因子:
13.6
通讯作者:
Nosek, Brian A.
Nosek, Brian A.
中科院分区:
心理学1区
文献类型:
--
作者:
Ebersole, Charles R.;Mathur, Maya B.;Nosek, Brian A.

文献摘要

被引文献

相似文献

心理科学中的重复研究有时无法重现先前的发现。如果这些研究使用的方法不忠实于原始研究或在引发感兴趣的现象方面无效,则未能复制可能是方案失败,而不是对原始发现的挑战。由专家进行正式的数据收集前同行审查,可解决不足之处,提高可复制率。我们从生殖项目:心理学(RP:P;开放科学合作,2015)中选择了10项复制研究,其原始作者在数据收集前表达了对复制设计的担忧;其中只有一项研究产生了统计学显著性影响(p < .05)。评论者认为,缺乏对专家评审的坚持和低功率测试是大多数RP:P研究未能复制原始效果的原因。我们修订了复制方案,并在进行新的复制研究之前接受了正式的同行评审。我们在多个实验室(每个原始研究的实验室中位数= 6.5,范围= 3-9;总样本中位数= 1,279.5,范围= 276- 3,512)中使用RP:P和修订后的方案,对两种方案的每个原始结果进行高倍检验。总体而言,根据预注册的分析计划,我们发现修订后的方案产生的效应量与RP:P方案相似(Delta r = .002或.014,取决于分析方法)。修订方案的中位效应量(r = 0.05)与RP:P方案(r = 0.04)和原始RP:P重复(r = 0.11)相似,小于原始研究(r = 0.37)。对原始研究和相应的三次重复尝试的累积证据的分析提供了对10种测试效应的非常精确的估计,并表明其效应量(中位数r = 0.07,范围= 0.00 - 0.15)平均比原始效应量(中位数r = 0.37,范围= 0.19 - 0.50)小78%。
Replication studies in psychological science sometimes fail to reproduce prior findings. If these studies use methods that are unfaithful to the original study or ineffective in eliciting the phenomenon of interest, then a failure to replicate may be a failure of the protocol rather than a challenge to the original finding. Formal pre-data-collection peer review by experts may address shortcomings and increase replicability rates. We selected 10 replication studies from the Reproducibility Project: Psychology (RP:P; Open Science Collaboration, 2015) for which the original authors had expressed concerns about the replication designs before data collection; only one of these studies had yielded a statistically significant effect (p < .05). Commenters suggested that lack of adherence to expert review and low-powered tests were the reasons that most of these RP:P studies failed to replicate the original effects. We revised the replication protocols and received formal peer review prior to conducting new replication studies. We administered the RP:P and revised protocols in multiple laboratories (median number of laboratories per original study = 6.5, range = 3-9; median total sample = 1,279.5, range = 276-3,512) for high-powered tests of each original finding with both protocols. Overall, following the preregistered analysis plan, we found that the revised protocols produced effect sizes similar to those of the RP:P protocols (Delta r = .002 or .014, depending on analytic approach). The median effect size for the revised protocols (r = .05) was similar to that of the RP:P protocols (r = .04) and the original RP:P replications (r = .11), and smaller than that of the original studies (r = .37). Analysis of the cumulative evidence across the original studies and the corresponding three replication attempts provided very precise estimates of the 10 tested effects and indicated that their effect sizes (median r = .07, range = .00-.15) were 78% smaller, on average, than the original effect sizes (median r = .37, range = .19-.50).