Robust feature screening procedures for single and mixed types of data

Robust feature screening procedures for single and mixed types of data
复制标题

DOI:
10.1080/00949655.2020.1719104
复制
发表时间:
2020-02
影响因子:
1.2
通讯作者:
Jinhui Sun;Pang Du;Hongyu Miao;Hua Liang
Jinhui Sun;Pang Du;Hongyu Miao;Hua Liang
中科院分区:
数学4区
文献类型:
--
作者:
Jinhui Sun;Pang Du;Hongyu Miao;Hua Liang

文献摘要

相似文献

摘要特征筛选过程的目的是降低具有指数增长维数的数据的维数。现有的程序都集中在一个单一类型的预测,这是所有连续或所有离散。它们不能处理混合类型的变量、离群值或非线性趋势。在本文中,我们首先提出了新的功能筛选程序(S)的不同的连续/离散组合的响应和预测变量。分别基于边缘斯皮尔曼相关、边缘方差分析、边缘Kruskal-Wallis检验、Kolmogorov-Smirnov检验、Mann-Whitney检验和平滑样条模型。进行了广泛的模拟研究,比较新的和现有的程序,确定一个最好的强大的筛选程序,为每一种类型的数据的目的。然后,我们将这些最佳筛选程序组合联合收割机,形成混合类型数据的鲁棒特征筛选程序。通过仿真研究和一个真实的例子,我们证明了它对异常值和模型误设定的鲁棒性。
ABSTRACT Feature screening procedures aim to reducing the dimensionality of data with exponentially-growing dimensions. Existing procedures all focused on a single type of predictors, which are either all continuous or all discrete. They cannot address mixed types of variables, outliers, or nonlinear trends. In this paper we first propose new feature screening procedure(s) for different continuous/discrete combinations of response and predictor variables. They are respectively based on marginal Spearman correlation, marginal ANOVA test, marginal Kruskal-Wallis test, Kolmogorov-Smirnov test, Mann-Whitney test, and smoothing splines modeling. Extensive simulation studies are performed to compare the new and existing procedures, with the aim of identifying a best robust screening procedure for each single type of data. Then we combine these best screening procedures to form the robust feature screening procedure for mixed type of data. We demonstrate its robustness against outliers and model misspecification through simulation studies and a real example.