l1-PENALIZED QUANTILE REGRESSION IN HIGH-DIMENSIONAL SPARSE MODELS

l1-PENALIZED QUANTILE REGRESSION IN HIGH-DIMENSIONAL SPARSE MODELS
复制标题

DOI:
10.1214/10-aos827
复制
发表时间:
2011-02-01
影响因子:
4.5
通讯作者:
Chernozhukov, Victor
Chernozhukov, Victor
中科院分区:
数学1区
文献类型:
--
作者:
Belloni, Alexandre;Chernozhukov, Victor

文献摘要

被引文献

相似文献

我们考虑中值回归,更普遍地是在高维稀疏模型中可能无限的分位数回归集合。在这些模型中,回归器p的数量非常大,可能比样本量N大,但是大多数S回归器仅对响应变量的每个条件分位数都有非零影响,其中S比N较慢。由于在这种情况下普通的分位数回归不一致,因此我们考虑l(1)二元分位数回归(L(1)-QR),该回归损失了回归系数的L(1) - 量子,以及后载后的。 QR估计器(l(1)-QR),将普通QR应用于L(1)-QR选择的模型。首先,我们表明在一般条件下l(1)-QR在近端速率下是一致的。 root s/n根log(p boolean或n),在分位数索引的(0,1)的紧凑型集合中均匀。在得出这一结果时,我们提出了部分关键的,数据驱动的罚款水平的选择,并表明它满足实现此速度的要求。其次,我们表明在类似条件下,l(1)-QR在近门速率root s/n root log(p boolean或n)上是一致的,即使l(1)-qr也均匀地在u上。 - 选择的模型错过了真实模型的某些组件,否则速率可能更接近Oracle速率。第三,我们表征了l(1)-QR包含真实模型作为子模型的条件,并在所选模型的维度上得出界限,均匀地在u上。我们还提供了条件,在这些条件下,势头限制选择最小的真实模型,均匀地超过u。
We consider median regression and, more generally, a possibly infinite collection of quantile regressions in high-dimensional sparse models. In these models, the number of regressors p is very large, possibly larger than the sample size n, but only at most s regressors have a nonzero impact on each conditional quantile of the response variable, where s grows more slowly than n. Since ordinary quantile regression is not consistent in this case, we consider l(1)-penalized quantile regression (l(1)-QR), which penalizes the l(1)-norm of regression coefficients, as well as the post-penalized QR estimator (post-l(1)-QR), which applies ordinary QR to the model selected by l(1)-QR. First, we show that under general conditions l(1)-QR is consistent at the near-oracle rate. root s/n root log(p boolean OR n), uniformly in the compact set u subset of (0, 1) of quantile indices. In deriving this result, we propose a partly pivotal, data-driven choice of the penalty level and show that it satisfies the requirements for achieving this rate. Second, we show that under similar conditions post-l(1)-QR is consistent at the near-oracle rate root s/n root log(p boolean OR n), uniformly over u, even if the l(1)-QR-selected models miss some components of the true models, and the rate could be even closer to the oracle rate otherwise. Third, we characterize conditions under which l(1)-QR contains the true model as a submodel, and derive bounds on the dimension of the selected model, uniformly over u; we also provide conditions under which hard-thresholding selects the minimal true model, uniformly over u.