Bayesian Tests to Quantify the Result of a Replication Attempt

Bayesian Tests to Quantify the Result of a Replication Attempt
复制标题

DOI:
10.1037/a0036731
复制
发表时间:
2014-08-01
影响因子:
4.1
通讯作者:
Wagenmakers, Eric-Jan
Wagenmakers, Eric-Jan
中科院分区:
心理学1区
文献类型:
--
作者:
Verhagen, Josine;Wagenmakers, Eric-Jan

文献摘要

被引文献

相似文献

复制尝试对于经验科学至关重要。成功的复制尝试会增加研究人员对效果存在的信心,而失败的复制尝试会引起怀疑和怀疑。然而,通常不清楚复制尝试在多大程度上会导致成功或失败。为了量化复制结果,我们提出了一种新颖的贝叶斯复制测试,该测试比较两个相互竞争的假设的充分性。第一个假设是怀疑论者的假设,认为该效应是虚假的;这是假设效应大小为零的零假设,H-0 : delta=0。第二个假设是支持者的假设,认为效果与原始研究中发现的效果一致,这种效果可以通过后验分布来量化。因此,第二个假设——复制假设——由 H-r : delta 给出,类似于“原始研究的后验分布”。 H-0 和 H-r 之间的加权似然比量化了数据为复制成功和失败提供的证据。除了新测试之外,我们还提出了其他几个贝叶斯测试,这些测试解决了与复制研究相关的不同但相关的问题。这些测试涉及单独实验的独立结论、原始实验和复制尝试之间效果大小的差异以及基于汇总结果的总体结论。总之,这套贝叶斯测试可以相对完整地形式化复制尝试的结果改变我们对当前现象的了解的方式。所有贝叶斯复制测试的使用均通过文献中的 3 个示例进行说明。对于使用 t 检验分析的实验,新复制检验的计算仅需要原始研究和复制研究中的 t 值和参与者数量。
Replication attempts are essential to the empirical sciences. Successful replication attempts increase researchers' confidence in the presence of an effect, whereas failed replication attempts induce skepticism and doubt. However, it is often unclear to what extent a replication attempt results in success or failure. To quantify replication outcomes we propose a novel Bayesian replication test that compares the adequacy of 2 competing hypotheses. The 1st hypothesis is that of the skeptic and holds that the effect is spurious; this is the null hypothesis that postulates a zero effect size, H-0 : delta=0. The 2nd hypothesis is that of the proponent and holds that the effect is consistent with the one found in the original study, an effect that can be quantified by a posterior distribution. Hence, the 2nd hypothesis-the replication hypothesis-is given by H-r : delta similar to "posterior distribution from original study." The weighted-likelihood ratio between H-0 and H-r quantifies the evidence that the data provide for replication success and failure. In addition to the new test, we present several other Bayesian tests that address different but related questions concerning a replication study. These tests pertain to the independent conclusions of the separate experiments, the difference in effect size between the original experiment and the replication attempt, and the overall conclusion based on the pooled results. Together, this suite of Bayesian tests allows a relatively complete formalization of the way in which the result of a replication attempt alters our knowledge of the phenomenon at hand. The use of all Bayesian replication tests is illustrated with 3 examples from the literature. For experiments analyzed using the t test, computation of the new replication test only requires the t values and the numbers of participants from the original study and the replication study.