Power for tests of interaction: effect of raising the Type I error rate.

Power for tests of interaction: effect of raising the Type I error rate.
复制标题

DOI:
10.1186/1742-5573-4-4
复制
发表时间:
2007-06-19
期刊:
Epidemiologic perspectives & innovations : EP+I
影响因子:
--
通讯作者:
Marshall, Stephen W
Marshall, Stephen W
中科院分区:
其他
文献类型:
--
作者:
Marshall, Stephen W

文献摘要

被引文献

相似文献

背景:在流行病学研究中,数据分析过程中评估相互作用的能力通常很差。这是因为流行病学研究通常主要只是为了评估主要影响。有鉴于此,一些研究人员在测试交互作用时提高了 I 类错误率,从而提高了功效。 However, this is a poor analysis strategy if the study is chronically under-powered (e.g. in a small study) or already adequately powered (e.g. in a very large study).为了证明这一点,本研究针对各种研究规模和交互类型,量化了当 I 类错误率升高时测试交互作用的功效增益。方法:计算交互作用的 Wald 检验、交互作用的似然比检验和比值比异质性的 Breslow-Day 检验的功效。在两个二元风险因素的简单场景中研究了从次加法到超乘法的十种类型的相互作用。研究了不同规模的病例对照研究(75 个病例和 150 个对照、300 个病例和 600 个对照、1200 个病例和 2400 个对照)。结果:将 I 类错误率从 5% 提高到 20% 的策略仅在所研究的 27 种交互类型/研究规模场景中的 7 种中产生了有用的功效增益(至少增益 10%,功效至少为 70%) (26%)。在其他 20 个场景中,功效要么已经足够(n = 8;30%),要么太低以至于即使将 I 类错误率提高到 20%(n = 12;44%),功效仍然很弱(低于 70%)。 结论:放宽 I 类错误率并不能有效地提高许多研究场景中交互测试的功效。在许多研究中,通过提高第一类误差所获得的小功率增益将被增加的“误报”的缺点所抵消。我建议研究者在评估交互测试时不应经常提高 I 类错误率。
BACKGROUND: Power for assessing interactions during data analysis is often poor in epidemiologic studies. This is because epidemiologic studies are frequently powered primarily to assess main effects only. In light of this, some investigators raise the Type I error rate, thereby increasing power, when testing interactions. However, this is a poor analysis strategy if the study is chronically under-powered (e.g. in a small study) or already adequately powered (e.g. in a very large study). To demonstrate this point, this study quantified the gain in power for testing interactions when the Type I error rate is raised, for a variety of study sizes and types of interaction.METHODS: Power was computed for the Wald test for interaction, the likelihood ratio test for interaction, and the Breslow-Day test for heterogeneity of the odds ratio. Ten types of interaction, ranging from sub-additive through to super-multiplicative, were investigated in the simple scenario of two binary risk factors. Case-control studies of various sizes were investigated (75 cases & 150 controls, 300 cases & 600 controls, and 1200 cases & 2400 controls).RESULTS: The strategy of raising the Type I error rate from 5% to 20% resulted in a useful power gain (a gain of at least 10%, resulting in power of at least 70%) in only 7 of the 27 interaction type/study size scenarios studied (26%). In the other 20 scenarios, power was either already adequate (n = 8; 30%), or else so low that it was still weak (below 70%) even after raising the Type I error rate to 20% (n = 12; 44%).CONCLUSION: Relaxing the Type I error rate did not usefully improve the power for tests of interaction in many of the scenarios studied. In many studies, the small power gains obtained by raising the Type I error will be more than offset by the disadvantage of increased "false positives". I recommend investigators should not routinely raise the Type I error rate when assessing tests of interaction.