From Fixed-X to Random-X Regression: Bias-Variance Decompositions, Covariance Penalties, and Prediction Error Estimation

From Fixed-X to Random-X Regression: Bias-Variance Decompositions, Covariance Penalties, and Prediction Error Estimation
复制标题

DOI:
10.1080/01621459.2018.1424632
复制
发表时间:
2020-01-02
影响因子:
3.7
通讯作者:
Tibshirani, Ryan J.
Tibshirani, Ryan J.
中科院分区:
数学1区
文献类型:
--
作者:
Rosset, Saharon;Tibshirani, Ryan J.

文献摘要

被引文献

相似文献

在统计预测中,基于协方差惩罚的模型选择和模型评估的经典方法仍然被广泛使用。关于这个主题的大多数文献都是基于我们所谓的“固定X”假设,其中协变量值被假设为非随机的。相比之下,采用“Random-X”视图通常更合理,其中协变量值独立绘制用于训练和预测。为了研究协方差惩罚在这种情况下的适用性,我们提出了一种随机X预测误差的分解,其中协变量中的随机性对偏差和方差分量都有贡献。这种分解是一般的,但我们集中在普通最小二乘(OLS)回归的基本情况。我们证明,在这种设置中,从固定X到随机X预测的移动会导致偏差和方差的增加。当协变量服从正态分布且线性模型无偏时,该分解中的所有项都是显式可计算的,这产生了Mallow ' Cp的扩展,我们称之为RCp。RCp也持有渐近某些类的非正态协变量。当噪声方差未知时,插入通常的无偏估计导致我们称之为与Sp密切相关的方法和广义交叉验证(GCV)。对于过度偏倚,我们提出了一个基于普通交叉验证(OCV)的“捷径公式”的估计,从而产生了一种我们称为RCp+的方法。理论论证和数值模拟表明,RCp+是典型的上级OCV,虽然差异很小。我们进一步研究其他流行的估计的随机X误差。我们对岭回归得到的令人惊讶的结果是,在高度正则化的情况下,Random-X方差小于Fixed-X方差,这可能导致更小的总体Random-X误差。本文的补充材料可在网上查阅。
In statistical prediction, classical approaches for model selection and model evaluation based on covariance penalties are still widely used. Most of the literature on this topic is based on what we call the "Fixed-X" assumption, where covariate values are assumed to be nonrandom. By contrast, it is often more reasonable to take a "Random-X" view, where the covariate values are independently drawn for both training and prediction. To study the applicability of covariance penalties in this setting, we propose a decomposition of Random-X prediction error in which the randomness in the covariates contributes to both the bias and variance components. This decomposition is general, but we concentrate on the fundamental case of ordinary least-squares (OLS) regression. We prove that in this setting the move from Fixed-X to Random-X prediction results in an increase in both bias and variance. When the covariates are normally distributed and the linear model is unbiased, all terms in this decomposition are explicitly computable, which yields an extension of Mallows' Cp that we call RCp. RCp also holds asymptotically for certain classes of nonnormal covariates. When the noise variance is unknown, plugging in the usual unbiased estimate leads to an approach that we call , which is closely related to Sp, and generalized cross-validation (GCV). For excess bias, we propose an estimate based on the "shortcut-formula" for ordinary cross-validation (OCV), resulting in an approach we call RCp+. Theoretical arguments and numerical simulations suggest that RCp+ is typically superior to OCV, though the difference is small. We further examine the Random-X error of other popular estimators. The surprising result we get for ridge regression is that, in the heavily regularized regime, Random-X variance is smaller than Fixed-X variance, which can lead to smaller overall Random-X error. Supplementary materials for this article are available online.