False discovery rate control via debiased lasso

False discovery rate control via debiased lasso
复制标题

DOI:
10.1214/19-ejs1554
复制
发表时间:
2019-01-01
影响因子:
1.1
通讯作者:
Javadi, Hamid
Javadi, Hamid
中科院分区:
数学3区
文献类型:
--
作者:
Javanmard, Adel;Javadi, Hamid

文献摘要

被引文献

相似文献

我们考虑高维统计模型中的变量选择问题,其中目标是从许多预测变量X-1,.,中报告一组变量与感兴趣的响应相关的X-p。对于线性高维模型,当参数个数超过样本个数(p > n)时,我们提出了一种变量选择方法,并证明了该方法可以控制方向错误发现率(FDR)低于预先指定的显著性水平q是[0,1]的元素.我们进一步分析了我们的框架的统计功效,并表明对于具有亚高斯行和公共精度矩阵的设计,如果最小非零参数θ(min)满足root n θ(min)- sigma root 2,则Omega是R-pxp的元素(最大Omega ii)(i是[p]的元素)log(2 p/qs(0))->无穷大,我们的框架建立在去偏方法的基础上,并假设标准条件s(0)= o(root n/(logp)(2)),其中s(0)指示p个特征中的真阳性的数量。值得注意的是,该框架实现了精确的方向FDR控制,而无需对未知回归参数的幅度进行任何假设,并且不需要任何协变量分布或噪声水平的知识。我们测试我们的方法在合成和真实的数据实验,以评估其性能,并证实我们的理论结果。
We consider the problem of variable selection in high-dimensional statistical models where the goal is to report a set of variables, out of many predictors X-1,...,X-p that are relevant to a response of interest. For linear high-dimensional model, where the number of parameters exceeds the number of samples (p > n), we propose a procedure for variables selection and prove that it controls the directional false discovery rate (FDR) below a pre-assigned significance level q is an element of [0, 1]. We further analyze the statistical power of our framework and show that for designs with subgaussian rows and a common precision matrix Omega is an element of R-pxp, if the minimum nonzero parameter theta(min) satisfiesroot n theta(min) - sigma root 2(max Omega ii)(i is an element of[p]) log(2p/qs(0)) -> infinity,then this procedure achieves asymptotic power one.Our framework is built upon the debiasing approach and assumes the standard condition s(0) = o(root n/(logp)(2)), where s(0) indicates the number of true positives among the p features. Notably, this framework achieves exact directional FDR control without any assumption on the amplitude of unknown regression parameters, and does not require any knowledge of the distribution of covariates or the noise level. We test our method in synthetic and real data experiments to assess its performance and to corroborate our theoretical results.