Simultaneous feature selection and outlier detection with optimality guarantees.

Simultaneous feature selection and outlier detection with optimality guarantees.
复制标题

DOI:
10.1111/biom.13553
复制
发表时间:
2022-12
期刊:
影响因子:
1.9
通讯作者:
Felici, Giovanni
Felici, Giovanni
中科院分区:
数学3区
文献类型:
--
作者:
Insolia, Luca;Kenney, Ana;Chiaromonte, Francesca;Felici, Giovanni

文献摘要

参考文献

相似文献

生物医学研究的数据越来越丰富,研究包括越来越多的功能。研究规模越大,大部分特征可能是冗余的和/或包含污染(离群值)的可能性就越高。这构成了严重的挑战,在样本规模相对较小的情况下,这种挑战更加严重。有效和高效的方法来执行稀疏估计存在离群值是这些研究的关键,并在过去十年中得到了相当大的关注。考虑到高维回归被影响响应和设计矩阵的多个均值漂移离群值所污染,我们对这一领域做出了贡献。我们开发了一个通用框架,并使用混合整数规划来同时执行特征选择和离群值检测,并具有可证明的最优保证。我们证明了我们的方法的理论性质,即,一个必要和充分条件的强大的预言家财产,其中的功能数量可以随样本量呈指数增长;参数的最佳估计;以及由此产生的估计的崩溃点。此外,我们提供了计算效率高的程序来调整整数约束和热启动算法。与现有的启发式方法相比,我们通过模拟展示了我们的建议的上级性能,并使用它来研究儿童肥胖和人类微生物组之间的关系。
Biomedical research is increasingly data rich, with studies comprising ever growing numbers of features. The larger a study, the higher the likelihood that a substantial portion of the features may be redundant and/or contain contamination (outlying values). This poses serious challenges, which are exacerbated in cases where the sample sizes are relatively small. Effective and efficient approaches to perform sparse estimation in the presence of outliers are critical for these studies, and have received considerable attention in the last decade. We contribute to this area considering high‐dimensional regressions contaminated by multiple mean‐shift outliers affecting both the response and the design matrix. We develop a general framework and use mixed‐integer programming to simultaneously perform feature selection and outlier detection with provably optimal guarantees. We prove theoretical properties for our approach, that is, a necessary and sufficient condition for the robustly strong oracle property, where the number of features can increase exponentially with the sample size; the optimal estimation of parameters; and the breakdown point of the resulting estimates. Moreover, we provide computationally efficient procedures to tune integer constraints and warm‐start the algorithm. We show the superior performance of our proposal compared to existing heuristic methods through simulations and use it to study the relationships between childhood obesity and the human microbiome.
DOI: 10.1080/00401706.1970.10488634
发表时间: 1970-01-01
期刊: TECHNOMETRICS
影响因子: 2.5
作者:
HOERL, AE;KENNARD, RW
通讯作者: KENNARD, RW
DOI: 10.1214/13-aos1198
发表时间: 2014-06
影响因子: 4.5
作者:
Fan J;Xue L;Zou H
通讯作者: Zou H
DOI: 10.1214/19-sts733
发表时间: 2020-11-01
影响因子: 5.7
作者:
Hastie, Trevor;Tibshirani, Robert;Tibshirani, Ryan
通讯作者: Tibshirani, Ryan
DOI: 10.1038/s41598-018-31866-9
发表时间: 2018-09-19
期刊: SCIENTIFIC REPORTS
影响因子: 4.6
作者:
Craig, Sarah J. C.;Blankenberg, Daniel;Makova, Kateryna D.
通讯作者: Makova, Kateryna D.
DOI: 10.1007/s10107-005-0594-3
发表时间: 2006-06-01
影响因子: 2.7
作者:
Frangioni, A;Gentile, C
通讯作者: Gentile, C