VARIABLE SELECTION FOR GENERAL INDEX MODELS VIA SLICED INVERSE REGRESSION

VARIABLE SELECTION FOR GENERAL INDEX MODELS VIA SLICED INVERSE REGRESSION
复制标题

DOI:
10.1214/14-aos1233
复制
发表时间:
2014-10-01
影响因子:
4.5
通讯作者:
Liu, Jun S.
Liu, Jun S.
中科院分区:
数学1区
文献类型:
--
作者:
Jiang, Bo;Liu, Jun S.

文献摘要

被引文献

相似文献

变量选择,也称为机器学习中的特征选择,在高维数据建模中起着重要作用,是数据驱动的科学发现的关键。我们在这里考虑的问题,检测有影响力的变量下的一般指数模型,其中的响应是依赖于预测通过一个未知的函数的一个或多个线性组合。而不是建立一个预测模型的预测组合的响应,我们模型的预测因子的条件分布给定的响应。这种逆向建模的角度促使我们提出了一个逐步的过程,基于似然比检验,这是有效的,计算效率高,在确定重要的变量,而不指定的预测和响应之间的参数关系。例如,所提出的过程能够检测p个预测变量之间具有成对、三向甚至更高阶交互的变量,计算时间为O(p)而不是O(p(k))(k是交互的最高阶)。通过仿真研究和真实的数据实例,证明了它与现有方法相比具有优异的经验性能。建立了当预测变量数和样本量都趋于无穷大时变量选择过程的一致性。
Variable selection, also known as feature selection in machine learning, plays an important role in modeling high dimensional data and is key to data-driven scientific discoveries. We consider here the problem of detecting influential variables under the general index model, in which the response is dependent of predictors through an unknown function of one or more linear combinations of them. Instead of building a predictive model of the response given combinations of predictors, we model the conditional distribution of predictors given the response. This inverse modeling perspective motivates us to propose a stepwise procedure based on likelihood-ratio tests, which is effective and computationally efficient in identifying important variables without specifying a parametric relationship between predictors and the response. For example, the proposed procedure is able to detect variables with pairwise, three-way or even higher-order interactions among p predictors with a computational time of O(p) instead of O(p(k)) (with k being the highest order of interactions). Its excellent empirical performance in comparison with existing methods is demonstrated through simulation studies as well as real data examples. Consistency of the variable selection procedure when both the number of predictors and the sample size go to infinity is established.