Testing for the presence of significant covariates through conditional marginal regression

Testing for the presence of significant covariates through conditional marginal regression
复制标题

DOI:
10.1093/biomet/asx061
复制
发表时间:
2018-03
期刊:
影响因子:
2.7
通讯作者:
Yanlin Tang;H. Wang;Emre Barut
Yanlin Tang;H. Wang;Emre Barut
中科院分区:
数学2区
文献类型:
--
作者:
Yanlin Tang;H. Wang;Emre Barut

文献摘要

被引文献

相似文献

研究人员有时对预测因子的相对重要性有一个先验信息,可以用来筛选协变量。一个重要的问题是,当最相关的预测因子包含在模型中时,是否有任何丢弃的协变量具有预测能力。我们考虑在一些预先选择的协变量上检验任何丢弃的协变量是否显著。我们提出了一个极大型检验统计量,并证明它具有非标准渐近分布,从而产生了条件自适应重抽样检验。为了适应未知稀疏度的信号,我们开发了一种混合检验统计量,它是最大型和和型统计量的加权平均值。我们在一般假设下证明了检验过程的一致性,并说明了如何将其用作正向回归的停止规则。我们通过仿真表明,即使在高维情况下,所提出的方法对稀疏和密集信号都具有竞争力的家族错误率提供了充分的控制,并且我们证明了它在协变量高度相关的情况下的优势。我们通过分析一个表达数量性状位点数据集来说明我们的方法的应用。
Summary Researchers sometimes have a priori information on the relative importance of predictors that can be used to screen out covariates. An important question is whether any of the discarded covariates have predictive power when the most relevant predictors are included in the model. We consider testing whether any discarded covariate is significant conditional on some pre-chosen covariates. We propose a maximum-type test statistic and show that it has a nonstandard asymptotic distribution, giving rise to the conditional adaptive resampling test. To accommodate signals of unknown sparsity, we develop a hybrid test statistic, which is a weighted average of maximum- and sum-type statistics. We prove the consistency of the test procedure under general assumptions and illustrate how it can be used as a stopping rule in forward regression. We show, through simulation, that the proposed method provides adequate control of the familywise error rate with competitive power for both sparse and dense signals, even in high-dimensional cases, and we demonstrate its advantages in cases where the covariates are heavily correlated. We illustrate the application of our method by analysing an expression quantitative trait locus dataset.