Robust variable selection and estimation via adaptive elastic net S-estimators for linear regression

Robust variable selection and estimation via adaptive elastic net S-estimators for linear regression
复制标题

DOI:
10.1016/j.csda.2023.107730
复制
发表时间:
2021-07
期刊:
Comput. Stat. Data Anal.
影响因子:
--
通讯作者:
D. Kepplinger
D. Kepplinger
中科院分区:
其他
文献类型:
--
作者:
D. Kepplinger

文献摘要

相似文献

具有异常值的重尾误差分布和预测因子在高维回归问题中普遍存在,如果处理不当,可能严重危及统计分析的有效性。为了在这些不利条件下更可靠地选择和预测变量,提出了一种新的鲁棒正则化回归估计器自适应PENSE。自适应PENSE产生可靠的变量选择和系数估计,即使在异常污染的预测或残差。结果表明,自适应惩罚比其他惩罚具有更强的鲁棒性和可靠性,特别是在预测空间中存在粗异常值的情况下。进一步证明了自适应PENSE具有较强的变量选择特性,即使在重尾误差下也具有oracle特性,无需估计误差尺度。对模拟和真实数据集的数值研究突出了在污染样本情况下,与其他鲁棒正则化估计器相比,在广泛的设置范围内具有优越的有限样本性能。在补充资料中提供了实现快速算法的R包和附加的仿真结果。
Heavy-tailed error distributions and predictors with anomalous values are ubiquitous in high-dimensional regression problems and can seriously jeopardize the validity of statistical analyses if not properly addressed. For more reliable variable selection and prediction under these adverse conditions, adaptive PENSE, a new robust regularized regression estimator, is proposed. Adaptive PENSE yields reliable variable selection and coefficient estimates even under aberrant contamination in the predictors or residuals. It is shown that the adaptive penalty leads to more robust and reliable variable selection than other penalties, particularly in the presence of gross outliers in the predictor space. It is further demonstrated that adaptive PENSE has strong variable selection properties and that it possesses the oracle property even under heavy-tailed errors and without the need to estimate the error scale. Numerical studies on simulated and real data sets highlight the superior finite-sample performance in a vast range of settings compared to other robust regularized estimators in the case of contaminated samples. An R package implementing a fast algorithm for computing the proposed method and additional simulation results are provided in the supplementary materials.