Confidence intervals and hypothesis testing for high-dimensional regression

Confidence intervals and hypothesis testing for high-dimensional regression
复制标题

DOI:
10.5555/2627435.2697057
复制
发表时间:
2013-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Adel Javanmard;A. Montanari
Adel Javanmard;A. Montanari
中科院分区:
其他
文献类型:
--
作者:
Adel Javanmard;A. Montanari

文献摘要

被引文献

相似文献

拟合高维统计模型通常需要使用非线性参数估计程序。因此,通常不可能获得参数估计的概率分布的精确表征。这反过来意味着量化与特定参数估计相关的不确定性极具挑战性。具体而言,不存在普遍接受的程序来计算不确定性和统计显着性的经典度量作为这些模型的置信区间或 p 值。我们在这里考虑高维线性回归问题,并提出一种用于构建置信区间和 p 值的有效算法。得到的置信区间具有接近最佳的大小。当测试某个参数消失的零假设时,我们的方法具有接近最佳的功效。我们的方法基于构建正则化 M 估计量的“去偏”版本。新结构比该领域最近的工作有所改进,因为它在设计矩阵上没有采用特殊的结构。我们在合成数据和有关核黄素生产率的高通量基因组数据集上测试了我们的方法,这些数据集由 Buhlmann 等人公开提供。 (2014)。
Fitting high-dimensional statistical models often requires the use of non-linear parameter estimation procedures. As a consequence, it is generally impossible to obtain an exact characterization of the probability distribution of the parameter estimates. This in turn implies that it is extremely challenging to quantify the uncertainty associated with a certain parameter estimate. Concretely, no commonly accepted procedure exists for computing classical measures of uncertainty and statistical significance as confidence intervals or p- values for these models. We consider here high-dimensional linear regression problem, and propose an efficient algorithm for constructing confidence intervals and p-values. The resulting confidence intervals have nearly optimal size. When testing for the null hypothesis that a certain parameter is vanishing, our method has nearly optimal power. Our approach is based on constructing a 'de-biased' version of regularized M-estimators. The new construction improves over recent work in the field in that it does not assume a special structure on the design matrix. We test our method on synthetic data and a high-throughput genomic data set about riboflavin production rate, made publicly available by Buhlmann et al. (2014).