A comment on replication, p‐values and evidence S.N.Goodman, Statistics in Medicine 1992; 11:875‐879
A comment on replication, p‐values and evidence S.N.Goodman, Statistics in Medicine 1992; 11:875‐879
复制标题
对复制、p 值和证据的评论 S.N.Goodman,《医学统计》1992 年 11:875‐879;
DOI:
10.1002/sim.1072
复制
发表时间:
2002
影响因子:
2
通讯作者:
S. Senn
中科院分区:
文献类型:
--
作者:
S. Senn
Some years ago, in the pages of this journal, Goodman gave an interesting analysis of ‘replication probabilities’ of p-values. Specifically, he considered the possibility that a given experiment had produced a p-value that indicated ‘significance’or near significance (he considered the range p= 0: 10 to 0.001) and then calculated the probability that a study with equal power would produce a significant result at the conventional level of significance of 0.05. He showed, for example, that given an uninformative prior, and (subsequently) a resulting p-value that was exactly 0.05 from the first experiment, the probability of significance in the second experiment was 50 per cent. A more general form of this result is as follows. If the first trial yields p= then the probability that a second trial will be significant at significance level(and in the same direction as the first trial) is 0.5. I share many of Goodman’s misgiving about p-values and I do not disagree with his calculations (except in slight numerical details). I also consider that his demonstration is useful for two reasons. First, it serves as a warning for anybody planning a further similar study to one just completed (and which has a marginally significant result) that this may not be matched in the second study. Second, it serves as a warning that apparent inconsistency in results from individual studies may be expected to be common and that one should not overreact to this phenomenon.However, I disagree with two points that he makes. First, he claims that ‘the replication probability provides a means, within the frequentist framework, to separate p-values from their hypothesis test interpretation, an important first step towards understanding the concept of inferential meaning’(p. 879). I disagree with him on two grounds here:(i) it is not necessary to separate p-values from their hypothesis test interpretation;(ii) the replication probability has no direct bearing on inferential meaning. Second he claims that,‘the replication probability can be used as a frequentist counterpart of Bayesian and likelihood models to show that p-values overstate the evidence against the null-hypothesis’(p. 875, my italics). I disagree that there is such an overstatement.