A comment on replication, p‐values and evidence S.N.Goodman, Statistics in Medicine 1992; 11:875‐879

A comment on replication, p‐values and evidence S.N.Goodman, Statistics in Medicine 1992; 11:875‐879
复制标题

对复制、p 值和证据的评论 S.N.Goodman,《医学统计》1992 年 11:875‐879;

DOI:
10.1002/sim.1072
复制
发表时间:
2002
影响因子:
2
通讯作者:
S. Senn
S. Senn
中科院分区:
医学3区
文献类型:
--
作者:
S. Senn

文献摘要

被引文献

相似文献

几年前,在这本杂志的页面上,古德曼对 p 值的“复制概率”进行了有趣的分析。具体来说,他考虑了给定实验产生的 p 值表明“显着性”或接近显着性的可能性(他考虑了范围 p= 0: 10 到 0.001),然后计算了具有相同功效的研究在 0.05 的传统显着性水平上产生显着性结果的概率。例如,他表明,如果先验信息不丰富,并且(随后)第一个实验的结果 p 值恰好为 0.05,则第二个实验的显着性概率为 50%。该结果的更一般形式如下。如果第一次试验的结果为 p=,则第二次试验在显着性水平上显着(并且与第一次试验的方向相同)的概率为 0.5。我和古德曼一样对 p 值抱有许多疑虑,并且我并不反对他的计算(除了一些细微的数字细节)。我还认为他的演示很有用,原因有两个。首先,它对任何计划进行与刚刚完成的一项进一步类似研究(并且具有轻微显着性结果)的人发出警告,第二项研究可能与第二项研究不匹配。其次,它是一个警告,表明个别研究结果的明显不一致可能是很常见的,人们不应该对这种现象反应过度。但是,我不同意他提出的两点。首先,他声称“在频率论框架内,复制概率提供了一种将 p 值与其假设检验解释分开的方法,这是理解推理意义概念的重要的第一步”(第 879 页)。我在这里有两个理由不同意他的观点:(i) 没有必要将 p 值与其假设检验解释分开;(ii) 复制概率对推理意义没有直接影响。其次,他声称,“复制概率可以用作贝叶斯模型和似然模型的频率论对应物,以表明 p 值夸大了反对零假设的证据”(第 875 页,我的斜体字)。我不同意这种夸张的说法。
Some years ago, in the pages of this journal, Goodman gave an interesting analysis of ‘replication probabilities’ of p-values. Specifically, he considered the possibility that a given experiment had produced a p-value that indicated ‘significance’or near significance (he considered the range p= 0: 10 to 0.001) and then calculated the probability that a study with equal power would produce a significant result at the conventional level of significance of 0.05. He showed, for example, that given an uninformative prior, and (subsequently) a resulting p-value that was exactly 0.05 from the first experiment, the probability of significance in the second experiment was 50 per cent. A more general form of this result is as follows. If the first trial yields p= then the probability that a second trial will be significant at significance level(and in the same direction as the first trial) is 0.5. I share many of Goodman’s misgiving about p-values and I do not disagree with his calculations (except in slight numerical details). I also consider that his demonstration is useful for two reasons. First, it serves as a warning for anybody planning a further similar study to one just completed (and which has a marginally significant result) that this may not be matched in the second study. Second, it serves as a warning that apparent inconsistency in results from individual studies may be expected to be common and that one should not overreact to this phenomenon.However, I disagree with two points that he makes. First, he claims that ‘the replication probability provides a means, within the frequentist framework, to separate p-values from their hypothesis test interpretation, an important first step towards understanding the concept of inferential meaning’(p. 879). I disagree with him on two grounds here:(i) it is not necessary to separate p-values from their hypothesis test interpretation;(ii) the replication probability has no direct bearing on inferential meaning. Second he claims that,‘the replication probability can be used as a frequentist counterpart of Bayesian and likelihood models to show that p-values overstate the evidence against the null-hypothesis’(p. 875, my italics). I disagree that there is such an overstatement.