SLICED INVERSE REGRESSION FOR DIMENSION REDUCTION

SLICED INVERSE REGRESSION FOR DIMENSION REDUCTION
复制标题

DOI:
10.2307/2290563
复制
发表时间:
1991-06-01
影响因子:
3.7
通讯作者:
LI, KC
LI, KC
中科院分区:
数学1区
文献类型:
--
作者:
LI, KC

文献摘要

被引文献

相似文献

计算能力的现代进步极大地拓宽了科学家从许多变量中收集和研究信息的范围,这些信息在过去可能被忽视。 然而,有效地扫描一个大的变量池并不是一件容易的事情,尽管我们与数据交互的能力已经通过动态图形的最新创新得到了很大的增强。 在这篇文章中,我们提出了一种新的数据分析工具,切片逆回归(SIR),用于减少输入变量x的维数,而不需要经过任何参数或非参数模型拟合过程。 这种方法探索了回归的逆视图的简单性;也就是说,不是回归单变量输出变量y对多变量X,而是回归x对y。 正向回归和反向回归由一个激励这种方法的定理连接。 在模型y = f(beta-1x,.,beta(K)x,其中beta-k是未知的行向量。 这个模型看起来像一个非线性回归,除了f的函数形式是完全未知的。 为了有效地降低维度,我们只需要估计空间[有效降维(e.d.r.)由beta-k生成的空间。 这使得我们的目标不同于回归分析中通常的目标,即所有回归系数的估计。 事实上,β-k本身在f上没有特定的结构形式是无法识别的。 我们的主要定理表明,在适当的条件下,如果x的分布被标准化为具有零均值和单位协方差,则逆回归曲线E(x \ y)将落入e.d.r.空间 因此,可以对估计的逆回归曲线的协方差矩阵进行主成分分析,以确定其主方向,从而得到我们的估计e.d.r.。方向 此外,我们使用一个简单的阶梯函数来估计逆回归曲线。 不需要复杂的平滑。 SIR可以很容易地在个人计算机上实现。 通过模拟,我们展示了SIR如何有效地降低输入变量的维数,比如说,从10到K = 2的数据集与400个观察。 由SIR得到的y对两个投影变量的自旋图很好地模拟了y对真实方向的自旋图。 卡方统计量提出了解决的问题,是否由SIR发现的方向是虚假的。
Modern advances in computing power have greatly widened scientists' scope in gathering and investigating information from many variables, information which might have been ignored in the past. Yet to effectively scan a large pool of variables is not an easy task, although our ability to interact with data has been much enhanced by recent innovations in dynamic graphics. In this article, we propose a novel data-analytic tool, sliced inverse regression (SIR), for reducing the dimension of the input variable x without going through any parametric or nonparametric model-fitting process. This method explores the simplicity of the inverse view of regression; that is, instead of regressing the univariate output variable y against the multivariate X, we regress x against y. Forward regression and inverse regression are connected by a theorem that motivates this method. The theoretical properties of SIR are investigated under a model of the form, y = f(beta-1x, ..., beta(K)x, epsilon), where the beta-k's are the unknown row vectors. This model looks like a nonlinear regression, except for the crucial difference that the functional form of f is completely unknown. For effectively reducing the dimension, we need only to estimate the space [effective dimension reduction (e.d.r.) space] generated by the beta-k's. This makes our goal different from the usual one in regression analysis, the estimation of all the regression coefficients. In fact, the beta-k's themselves are not identifiable without a specific structural form on f. Our main theorem shows that under a suitable condition, if the distribution of x has been standardized to have the zero mean and the identity covariance, the inverse regression curve, E(x \ y), will fall into the e.d.r. space. Hence a principal component analysis on the covariance matrix for the estimated inverse regression curve can be conducted to locate its main orientation, yielding our estimates for e.d.r. directions. Furthermore, we use a simple step function to estimate the inverse regression curve. No complicated smoothing is needed. SIR can be easily implemented on personal computers. By simulation, we demonstrate how SIR can effectively reduce the dimension of the input variable from, say, 10 to K = 2 for a data set with 400 observations. The spin-plot of y against the two projected variables obtained by SIR is found to mimic the spin-plot of y against the true directions very well. A chi-squared statistic is proposed to address the issue of whether or not a direction found by SIR is spurious.