Ranking-Based Variable Selection for high-dimensional data

Ranking-Based Variable Selection for high-dimensional data
复制标题

DOI:
10.5705/ss.202017.0139
复制
发表时间:
2020
期刊:
影响因子:
1.4
通讯作者:
R. Baranowski;Yining Chen;P. Fryzlewicz
R. Baranowski;Yining Chen;P. Fryzlewicz
中科院分区:
数学3区
文献类型:
--
作者:
R. Baranowski;Yining Chen;P. Fryzlewicz

文献摘要

相似文献

我们提出了一种基于排名的变量选择(RBVS)技术,可以识别影响高维数据响应的重要变量。 RBVS 使用子采样来识别非虚假出现在所选变量排名顶部的协变量。我们研究了这样一个集合是唯一的条件,并表明它可以通过我们的程序从数据中成功恢复。与许多现有的高维变量选择技术不同,RBVS 在所有相关变量中区分重要变量和不重要变量,并旨在仅恢复重要变量。此外,RBVS 不需要对响应和协变量之间的关系进行模型限制,因此广泛适用于参数和非参数环境。最后,我们通过比较仿真研究说明了所提出技术的良好实用性能。 RBVS 算法在 rbvs(一个公开可用的 R 包)中实现。
We propose a ranking-based variable selection (RBVS) technique that identifies important variables influencing the response in high-dimensional data. RBVS uses subsampling to identify the covariates that appear nonspuriously at the top of a chosen variable ranking. We study the conditions under which such a set is unique, and show that it can be recovered successfully from the data by our procedure. Unlike many existing high-dimensional variable selection techniques, among all relevant variables, RBVS distinguishes between important and unimportant variables, and aims to recover only the important ones. Moreover, RBVS does not require model restrictions on the relationship between the response and the covariates, and, thus, is widely applicable in both parametric and nonparametric contexts. Lastly, we illustrate the good practical performance of the proposed technique by means of a comparative simulation study. The RBVS algorithm is implemented in rbvs, a publicly available R package.