STATISTICAL METHODS FOR REPLICABILITY ASSESSMENT

STATISTICAL METHODS FOR REPLICABILITY ASSESSMENT
复制标题

DOI:
10.1214/20-aoas1336
复制
发表时间:
2020-09-01
影响因子:
1.8
通讯作者:
Fithian, William
Fithian, William
中科院分区:
数学4区
文献类型:
--
作者:
Hung, Kenneth;Fithian, William

文献摘要

被引文献

相似文献

大规模的复制研究,如复制计划:心理学(RP:P)提供了关于科学可复制性的宝贵系统数据,但大多数对数据的分析和解释未能就“可复制性”的定义达成一致,也未能将已知的选择偏差的必然后果与竞争性解释分开。我们讨论了可复制性的三个具体定义:(1)已发表的关于效应迹象的发现是否大多正确,(2)复制研究在复制原始实验中存在的真实效应大小方面的有效性如何,以及(3)真实效应大小是否倾向于在复制中减少。我们应用多重测试和后选择推理的技术来开发新的方法来回答这些问题,同时明确考虑选择偏差。我们的分析表明,RP:P数据集在很大程度上与由于选择显著效应而导致的发表偏倚一致。本文中的方法对真实效应量没有分布假设。
Large-scale replication studies like the Reproducibility Project: Psychology (RP:P) provide invaluable systematic data on scientific replicability, but most analyses and interpretations of the data fail to agree on the definition of "replicability" and disentangle the inexorable consequences of known selection bias from competing explanations. We discuss three concrete definitions of replicability based on: (1) whether published findings about the signs of effects are mostly correct, (2) how effective replication studies are in reproducing whatever true effect size was present in the original experiment and (3) whether true effect sizes tend to diminish in replication. We apply techniques from multiple testing and postselection inference to develop new methods that answer these questions while explicitly accounting for selection bias. Our analyses suggest that the RP:P dataset is largely consistent with publication bias due to selection of significant effects. The methods in this paper make no distributional assumptions about the true effect sizes.