Confronting p-hacking: addressing p-value dependence on sample size

Confronting p-hacking: addressing p-value dependence on sample size
复制标题

对抗 p-hacking:解决 p 值对样本大小的依赖性

DOI:
--
复制
发表时间:
2019
期刊:
bioRxiv
影响因子:
--
通讯作者:
A. Muñoz
A. Muñoz
中科院分区:
--
文献类型:
--
作者:
Estibaliz Gómez;A. Sneider;Hasini Jayatilaka;J. Phillip;D. Wirtz;A. Muñoz

文献摘要

被引文献

相似文献

生物医学研究已经开始依赖 p 值来确定潜在的转化影响。通常将 p 值与通常设置为 0.05 的阈值进行比较,以评估原假设的显着性。只要有足够大的数据集可用,就可以轻松达到此阈值。这种现象被称为 p-hacking,它会导致虚假结论。在此,我们提出了一种系统且易于遵循的协议,将 p 值建模为指数函数,以测试真实统计显着性的存在。这种新方法提供了对原假设的稳健评估,并提供了拒绝原假设所需的最小数据大小的准确值。通过模拟和实验获得的数据对该模型进行了深入研究。模拟表明,在受控数据下,我们的假设是正确的。我们对实验数据集的分析结果反映了这种方法在常见决策过程中的广泛应用。
Biomedical research has come to rely on p-values to determine potential translational impact. The p-value is routinely compared with a threshold commonly set to 0.05 to assess the significance of the null hypothesis. Whenever a large enough dataset is available, this threshold is easily reachable. This phenomenon is known as p-hacking and it leads to spurious conclusions. Herein, we propose a systematic and easy-to-follow protocol that models the p-value as an exponential function to test the existence of real statistical significance. This new approach provides a robust assessment of the null hypothesis with accurate values for the minimum data-size needed to reject it. An in-depth study of the model is carried out in both simulated and experimentally-obtained data. Simulations show that under controlled data, our assumptions are true. The results of our analysis in the experimental datasets reflect the large scope of this approach in common decision-making processes.