On the optimality of sliced inverse regression in high dimensions

On the optimality of sliced inverse regression in high dimensions
复制标题

DOI:
10.1214/19-aos1813
复制
发表时间:
2017-01
期刊:
The Annals of Statistics
影响因子:
--
通讯作者:
Q. Lin;Xinran Li;Dongming Huang;Jun S. Liu
Q. Lin;Xinran Li;Dongming Huang;Jun S. Liu
中科院分区:
其他
文献类型:
--
作者:
Q. Lin;Xinran Li;Dongming Huang;Jun S. Liu

文献摘要

被引文献

相似文献

随机变量对$(y,x)在{R}^{p+1}$中的中心子空间是使得$y\perp\hspace{-2 mm}\perp x\midP_{\mathcal{S}}x$的极小子空间。本文考虑多指标模型$y=f(β_{1}^{\tau}x,\beta_{2}^{\tau}x,...,β_d}^{\tau}x,\epsilon)$的中心空间估计的极小极大速度,其中$x\sim N(0,i_(P))$至多为$S有效预测元。我们首先引入了一大类依赖于$var(\mathbb{E}[x|y])$的最小非零本征值$\lambda$的模型,证明了基于SIR过程的聚集估计以$d\wedge((Sd+S\log(EP/S))/(n\lambda))$的速率收敛。然后,我们证明了在两种情况下,这个比率是最优的:单指数模型和中心维度固定的多指数模型和固定的$\lambda$。通过假设一个技术猜想,我们可以证明,对于中心空间维度有界的多指标模型,这个比率也是最优的。我们相信,这些(有条件的)最优利率结果给我们带来了对高维一般SDR问题的有意义的见解。
The central subspace of a pair of random variables $(y,x) \in \mathbb{R}^{p+1}$ is the minimal subspace $\mathcal{S}$ such that $y \perp \hspace{-2mm} \perp x\mid P_{\mathcal{S}}x$. In this paper, we consider the minimax rate of estimating the central space of the multiple index models $y=f(\beta_{1}^{\tau}x,\beta_{2}^{\tau}x,...,\beta_{d}^{\tau}x,\epsilon)$ with at most $s$ active predictors where $x \sim N(0,I_{p})$. We first introduce a large class of models depending on the smallest non-zero eigenvalue $\lambda$ of $var(\mathbb{E}[x|y])$, over which we show that an aggregated estimator based on the SIR procedure converges at rate $d\wedge((sd+s\log(ep/s))/(n\lambda))$. We then show that this rate is optimal in two scenarios: the single index models; and the multiple index models with fixed central dimension $d$ and fixed $\lambda$. By assuming a technical conjecture, we can show that this rate is also optimal for multiple index models with bounded dimension of the central space. We believe that these (conditional) optimal rate results bring us meaningful insights of general SDR problems in high dimensions.