Interpreting observational studies: why empirical calibration is needed to correct p-values.

Interpreting observational studies: why empirical calibration is needed to correct p-values.
复制标题

DOI:
10.1002/sim.5925
复制
发表时间:
2014-01-30
影响因子:
2
通讯作者:
Madigan, David
Madigan, David
中科院分区:
医学3区
文献类型:
--
作者:
Schuemie, Martijn J.;Ryan, Patrick B.;DuMouchel, William;Suchard, Marc A.;Madigan, David

文献摘要

参考文献

被引文献

相似文献

通常,文献根据‘p < 0.05’来断言医疗产品的效果。基本的前提是,在这个阈值下,只有5%的概率观察到的影响将被偶然看到,而实际上没有任何影响。在观察性研究中,与随机试验相比,偏见和混淆可能会破坏这一前提。为了验证这一前提,我们从文献中选择了三项样本药物安全性研究,分别代表病例对照、队列和自身对照病例系列设计。我们试图对原始文章中研究的药物尽可能地复制这些研究。接下来,我们将同样的三个设计应用于一组阴性对照:不被认为会导致感兴趣的结果的药物。我们观察了当零假设为真时p < 0.05时的频率,并将分布与效应估计值进行了拟合。利用这些分布,我们计算了反映在零假设下观察到效果估计的概率的校准p值,同时考虑了随机和系统误差。对科学文献进行了自动分析,以评估这种校准的潜在影响。我们的实验提供了证据,证明大多数观察性研究在没有影响的情况下会宣布统计意义。经验校准被发现可以将虚假结果减少到所需的5%的水平。将这些调整应用于文献表明,使用p < 0.05时,至少有54%的发现实际上没有统计学意义,应该重新评估。©2013作者。约翰·威利父子有限公司出版的医学统计数据。
Often the literature makes assertions of medical product effects on the basis of ‘ p < 0.05’. The underlying premise is that at this threshold, there is only a 5% probability that the observed effect would be seen by chance when in reality there is no effect. In observational studies, much more than in randomized trials, bias and confounding may undermine this premise. To test this premise, we selected three exemplar drug safety studies from literature, representing a case–control, a cohort, and a self-controlled case series design. We attempted to replicate these studies as best we could for the drugs studied in the original articles. Next, we applied the same three designs to sets of negative controls: drugs that are not believed to cause the outcome of interest. We observed how often p < 0.05 when the null hypothesis is true, and we fitted distributions to the effect estimates. Using these distributions, we compute calibrated p-values that reflect the probability of observing the effect estimate under the null hypothesis, taking both random and systematic error into account. An automated analysis of scientific literature was performed to evaluate the potential impact of such a calibration. Our experiment provides evidence that the majority of observational studies would declare statistical significance when no effect is present. Empirical calibration was found to reduce spurious results to the desired 5% level. Applying these adjustments to literature suggests that at least 54% of findings with p < 0.05 are not actually statistically significant and should be reevaluated. © 2013 The Authors. Statistics in Medicine published by John Wiley & Sons Ltd.
DOI: 10.1097/ede.0b013e3181d61eeb
发表时间: 2010-05
期刊: Epidemiology (Cambridge, Mass.)
影响因子: --
作者:
Lipsitch M;Tchetgen Tchetgen E;Cohen T
通讯作者: Cohen T
DOI: 10.1001/jama.2010.1098
发表时间: 2010-08-11
期刊: JAMA
影响因子: --
作者:
Cardwell CR;Abnet CC;Cantwell MM;Murray LJ
通讯作者: Murray LJ
DOI: 10.1111/j.1365-2036.2006.02985.x
发表时间: 2006-07-15
影响因子: 7.6
作者:
Abraham, NS;Cohen, DC;Richardson, P
通讯作者: Richardson, P
DOI: 10.1111/j.1572-0241.2007.01456.x
发表时间: 2007-11-01
影响因子: 9.8
作者:
Jinjuvadia, Kartik;Kwan, Wendy;Fontana, Robert J.
通讯作者: Fontana, Robert J.
DOI: 10.1371/journal.pmed.0020124
发表时间: 2005-08-01
期刊: PLOS MEDICINE
影响因子: 15.8
作者:
Ioannidis, JPA
通讯作者: Ioannidis, JPA