Subgroup analyses in randomized trials: risks of subgroup-specific analyses; power and sample size for the interaction test

Subgroup analyses in randomized trials: risks of subgroup-specific analyses; power and sample size for the interaction test
复制标题

DOI:
10.1016/j.jclinepi.2003.08.009
复制
发表时间:
2004-03-01
影响因子:
7.2
通讯作者:
Peters, TJ
Peters, TJ
中科院分区:
医学2区
文献类型:
--
作者:
Brookes, ST;Whitely, E;Peters, TJ

文献摘要

被引文献

相似文献

目的:尽管指南建议在临床试验的亚组分析中使用正式的相互作用检验,但不适当的亚组特异性分析仍在继续。此外,旨在检测总体治疗效果的试验检测治疗-亚组相互作用的能力有限。This article quantifies the error rate associated with subgroup analyses.Study Design and Setting:Simulations quantified the risk of misinterpreting subgroup analysis as evidence of different subgroup effects and the limited power of interaction test in trials designed to detect overall treatment effects.Results:虽然正式的相互作用检验在假阳性方面表现出了预期的效果,但亚组特异性检验的可靠性要低得多:根据试验特征,在7%至64%的模拟中仅观察到一个亚组的显著效应。关于交互作用检验的把握度,总体效应把握度为80%的试验仅具有29%的把握度来检测相同量级的交互作用。为了使这种大小的相互作用以与总体效应相同的功效被检测到,样本大小应该扩大四倍。结论:虽然人们普遍认为亚组分析可能产生虚假的结果,但问题的严重程度可能被低估。(C)2004年爱思唯尔公司All rights reserved.
Objective: Despite guidelines recommending the use of formal tests of interaction in subgroup analyses in clinical trials, inappropriate subgroup-specific analyses continue. Moreover, trials designed to detect overall treatment effects have limited power to detect treatment-subgroup interactions. This article quantifies the error rates associated with subgroup analyses.Study Design and Setting: Simulations quantified the risks of misinterpreting subgroup analyses as evidence of differential subgroup effects and the limited power of the interaction test in trials designed to detect overall treatment effects.Results: Although formal interaction tests performed as expected with respect to false positives, subgroup-specific tests were considerably less reliable: A significant effect in one subgroup only was observed in 7% to 64% of simulations depending on trial characteristics. Regarding power of the interaction test, a trial with 80% power for the overall effect had only 29% power to detect an interaction effect of the same magnitude. For interactions of this size to be detected with the same power as the overall effect, sample sizes should be inflated fourfold. increasing dramatically for interactions smaller than 20% of the overall effect.Conclusion: Although it is generally recognized that subgroup analyses can produce spurious results, the extent of the problem may be underestimated. (C) 2004 Elsevier Inc. All rights reserved.