Projective inference in high-dimensional problems: Prediction and feature selection

Projective inference in high-dimensional problems: Prediction and feature selection
复制标题

DOI:
10.1214/20-ejs1711
复制
发表时间:
2020-01-01
影响因子:
1.1
通讯作者:
Vehtari, Aki
Vehtari, Aki
中科院分区:
数学3区
文献类型:
--
作者:
Piironen, Juho;Paasiniemi, Markus;Vehtari, Aki

文献摘要

被引文献

相似文献

本文综述了高维数据稀缺的广义线性模型的预测推理和特征选择。我们证明,在许多情况下,人们可以从理论上合理的决策两阶段方法中受益:首先,构建一个可能的非稀疏模型,该模型可以很好地预测,然后找到表征预测的最小特征子集。第一步建立的模型称为参考模型,后一步的操作称为预测投影。这种方法的关键特征是它在稀疏性和预测准确性之间找到了一个很好的权衡,并且增益来自利用所有可用信息,包括先验信息和来自被遗漏的特征的信息。我们回顾了遵循这一原则的几种方法,并提供了新的方法贡献。我们提出了一种新的投影技术,将两种现有的投影技术结合起来,计算速度快,精度高。我们还提出了一种使用快速留一交叉验证来评估特征选择过程的方法,该方法允许简单直观的模型大小选择。此外,我们证明了一个定理,该定理有助于理解投影方法可能有益的条件。主要思想是通过几个实验,使用模拟和现实世界的数据说明。
This paper reviews predictive inference and feature selection for generalized linear models with scarce but high-dimensional data. We demonstrate that in many cases one can benefit from a decision theoretically justified two-stage approach: first, construct a possibly non-sparse model that predicts well, and then find a minimal subset of features that characterize the predictions. The model built in the first step is referred to as the reference model and the operation during the latter step as predictive projection. The key characteristic of this approach is that it finds an excellent tradeoff between sparsity and predictive accuracy, and the gain comes from utilizing all available information including prior and that coming from the left out features. We review several methods that follow this principle and provide novel methodological contributions. We present a new projection technique that unifies two existing techniques and is both accurate and fast to compute. We also propose a way of evaluating the feature selection process using fast leave-one-out cross-validation that allows for easy and intuitive model size selection. Furthermore, we prove a theorem that helps to understand the conditions under which the projective approach could be beneficial. The key ideas are illustrated via several experiments using simulated and real world data.