ROBUST RANK CORRELATION BASED SCREENING

ROBUST RANK CORRELATION BASED SCREENING
复制标题

基于稳健等级相关的筛选

DOI:
10.1214/12-aos1024
复制
发表时间:
2012-06-01
影响因子:
4.5
通讯作者:
Zhu, Lixing
Zhu, Lixing
中科院分区:
数学1区
文献类型:
--
作者:
Li, Gaorong;Peng, Heng;Zhu, Lixing

文献摘要

被引文献

相似文献

独立筛选是一种可变选择方法,它使用排名标准来选择重要变量,尤其是对于具有非多物质维度的统计模型或“大P,小N”范式的统计模型,当p可以像样本大小n的指数一样大。在本文中,我们提出了一种强大的等级相关筛选(RRCS)方法来处理超高维数据。新过程基于响应变量和预测变量之间的Kendall Tau相关系数,而不是现有方法的Pearson相关性。与现有的独立筛选方法相比,新方法具有四个理想的功能。首先,即使预测变量的数量在样本量的指数上呈指数的速度,即使预测变量的二阶矩矩的二阶矩,而不是指数式尾巴或雅致。其次,它可用于处理半参数模型,例如对链路函数的单调约束下的转换回归模型和单点数模型,即使模型中有非参数函数,也不涉及非参数估计。第三,该过程可以在很大程度上用于与异常值的使用,并影响观测值中的点。最后,与先前关于可变筛选的研究相比,由于所得统计数据的界限,指标函数在等级相关筛选中的使用极大地简化了理论推导。进行模拟以与现有方法进行比较,并分析一个真实的数据示例。
Independence screening is a variable selection method that uses a ranking criterion to select significant variables, particularly for statistical models with nonpolynomial dimensionality or "large p, small n" paradigms when p can be as large as an exponential of the sample size n. In this paper we propose a robust rank correlation screening (RRCS) method to deal with ultra-high dimensional data. The new procedure is based on the Kendall tau correlation coefficient between response and predictor variables rather than the Pearson correlation of existing methods. The new method has four desirable features compared with existing independence screening methods. First, the sure independence screening property can hold only under the existence of a second order moment of predictor variables, rather than exponential tails or alikeness, even when the number of predictor variables grows as fast as exponentially of the sample size. Second, it can be used to deal with semiparametric models such as transformation regression models and single-index models under monotonic constraint to the link function without involving nonparametric estimation even when there are nonparametric functions in the models. Third, the procedure can be largely used against outliers and influence points in the observations. Last, the use of indicator functions in rank correlation screening greatly simplifies the theoretical derivation due to the boundedness of the resulting statistics, compared with previous studies on variable screening. Simulations are carried out for comparisons with existing methods and a real data example is analyzed.