STATISTICAL POWER, SAMPLE-SIZE, AND THEIR REPORTING IN RANDOMIZED CONTROLLED TRIALS

STATISTICAL POWER, SAMPLE-SIZE, AND THEIR REPORTING IN RANDOMIZED CONTROLLED TRIALS
复制标题

DOI:
10.1001/jama.272.2.122
复制
发表时间:
1994-07-13
影响因子:
120.7
通讯作者:
WELLS, GA
WELLS, GA
中科院分区:
医学1区
文献类型:
--
作者:
MOHER, D;DULBERG, CS;WELLS, GA

文献摘要

被引文献

相似文献

目的。在统计能力级别中描述随着时间的流逝的模式,并在已发表的随机对照试验(RCT)中报告样本量计算,并取得负面结果。Desesign.Desesign.-Dessign.-我们的研究是一项描述性的调查。计算了检测25%和50%相对差异的功率,该子集使用了负面结果,其中使用了简单的两组并联设计。制定了标准以将试验结果归类为正面或负面,并确定主要结果。电源计算基于试验中报道的主要结果的结果。我们审查了1975年,1980年,1985年和1990年在JAMA,Lancet和New England Medicine杂志上发表的所有383个RCT。 - 在383个RCT(n = 102)中,有7%分类为负结果。从1975年到1990年,已发表的RCT数量增加了一倍以上,结果的比例保持不变。在具有二分法或连续原发性结果负面结果的简单两组并行设计试验中(n = 70),只有16%和36%具有足够的统计能力(80%),以分别检测到相对差异25%或50% 。这些百分比并没有始终增加加班。总体而言,只有32%的试验结果报告了样本量计算,但是随着时间的推移,1975年的0%提高到了1990年的43%。只有20个报告中只有20个报告发表了与临床意义有关的任何陈述在观察到的差异中,有否定结果的最大试验没有足够大的样本量来检测25%或50%的相对差异。随着时间的流逝,此结果并没有改变。很少有试验讨论观察到的差异在临床上是否重要。有重要的理由改变这种做法。统计功率和样本量的报告也需要改进。
Objective.-To describe the pattern over time in the level of statistical power and the reporting of sample size calculations in published randomized controlled trials (RCTs) with negative results.Design.-Our study was a descriptive survey. Power to detect 25% and 50% relative differences was calculated for the subset of trials with negative results in which a simple two-group parallel design was used. Criteria were developed both to classify trial results as positive or negative and to identify the primary outcomes. Power calculations were based on results from the primary outcomes reported in the trials.Population.-We reviewed all 383 RCTs published in JAMA, Lancet, and the New England Journal of Medicine in 1975, 1980, 1985, and 1990.Results.-Twenty-seven percent of the 383 RCTs (n=102) were classified as having negative results. The number of published RCTs more than doubled from 1975 to 1990, with the proportion of trials with negative results remaining fairly stable. Of the simple two-group parallel design trials having negative results with dichotomous or continuous primary outcomes (n=70), only 16% and 36% had sufficient statistical power (80%) to detect a 25% or 50% relative difference, respectively. These percentages did not consistently increase overtime. Overall, only 32% of the trials with negative results reported sample size calculations, but the percentage doing so has improved over time from 0% in 1975 to 43% in 1990. Only 20 of the 102 reports made any statement related to the clinical significance of the observed differences.Conclusions.-Most trials with negative results did not have large enough sample sizes to detect a 25% or a 50% relative difference. This result has not changed over time. Few trials discussed whether the observed differences were clinically important. There are important reasons to change this practice. The reporting of statistical power and sample size also needs to be improved.