Interaction analysis under misspecification of main effects: Some common mistakes and simple solutions.

Interaction analysis under misspecification of main effects: Some common mistakes and simple solutions.
复制标题

主效应错误指定下的交互分析:一些常见错误和简单的解决方案。

DOI:
10.1002/sim.8505
复制
发表时间:
2020
影响因子:
2
通讯作者:
Mukherjee,Bhramar
Mukherjee,Bhramar
中科院分区:
医学3区
文献类型:
--
作者:
Zhang,Min;Yu,Youfei;Wang,Shikun;Salvatore,Maxwell;GFritsche,Lars;He,Zihuai;Mukherjee,Bhramar

文献摘要

相似文献

在统计学和流行病学文献中,用两个线性主效应和一个产品项来模拟交互作用的统计实践是普遍存在的。大多数数据建模师都知道,主效应的错误指定可能会在交互测试中导致严重的I类错误膨胀,从而导致交互的错误检测。然而,建模实践并没有改变。在这篇文章中,我们专注于模型中的主要效应被错误指定为线性项的特定情况,并表征其对统计交互作用的常见检验的影响。然后,我们提出了一些简单的替代方案,以解决由于主效应错误指定而导致的测试交互中潜在的I型错误膨胀问题。我们证明了当对一个具有定量结果和两个独立因素的线性回归模型使用夹心方差估计时,Wald检验和SCORE检验都渐近地保持了正确的I型错误率。然而,如果独立性假设不成立,或者结果是二元的,那么使用三明治估计并不能解决问题。我们进一步证明,在广义加性模型下灵活地对主效应进行建模可以在很大程度上减少或经常消除估计中的偏差,并保持定量和二元结果的正确的I类错误率,而不考虑独立性假设。我们表明,在独立性假设下,对于连续的结果,相对于正确指定的主效应模型,过拟合和灵活地建模主效应并不会导致功率损失。我们的模拟研究进一步证明,使用灵活的主效应模型并不会导致总体上测试交互作用的能力显著下降。我们的结果提供了一个更好的理解,在存在主效应错误指定的情况下,交互作用测试的优势和局限性。使用来自大型生物库研究“密歇根基因组学倡议”的数据,我们提供了两个相互作用分析的例子来支持我们的结果。
The statistical practice of modeling interaction with two linear main effects and a product term is ubiquitous in the statistical and epidemiological literature. Most data modelers are aware that the misspecification of main effects can potentially cause severe type I error inflation in tests for interactions, leading to spurious detection of interactions. However, modeling practice has not changed. In this article, we focus on the specific situation where the main effects in the model are misspecified as linear terms and characterize its impact on common tests for statistical interaction. We then propose some simple alternatives that fix the issue of potential type I error inflation in testing interaction due to main effect misspecification. We show that when using the sandwich variance estimator for a linear regression model with a quantitative outcome and two independent factors, both the Wald and score tests asymptotically maintain the correct type I error rate. However, if the independence assumption does not hold or the outcome is binary, using the sandwich estimator does not fix the problem. We further demonstrate that flexibly modeling the main effect under a generalized additive model can largely reduce or often remove bias in the estimates and maintain the correct type I error rate for both quantitative and binary outcomes regardless of the independence assumption. We show, under the independence assumption and for a continuous outcome, overfitting and flexibly modeling the main effects does not lead to power loss asymptotically relative to a correctly specified main effect model. Our simulation study further demonstrates the empirical fact that using flexible models for the main effects does not result in a significant loss of power for testing interaction in general. Our results provide an improved understanding of the strengths and limitations for tests of interaction in the presence of main effect misspecification. Using data from a large biobank study “The Michigan Genomics Initiative”, we present two examples of interaction analysis in support of our results.