A scalable estimate of the out-of-sample prediction error via approximate leave-one-out cross-validation

A scalable estimate of the out-of-sample prediction error via approximate leave-one-out cross-validation
复制标题

DOI:
10.1111/rssb.12374
复制
发表时间:
2020-06-20
影响因子:
5.8
通讯作者:
Maleki, Arian
Maleki, Arian
中科院分区:
数学1区
文献类型:
--
作者:
Rad, Kamiar Rahnama;Maleki, Arian

文献摘要

被引文献

相似文献

本文研究了高维环境下样本外风险估计问题,其中标准技术如K折交叉验证存在较大偏差。受留一法交叉验证方法的低偏差的启发,我们提出了一个计算效率高的封闭形式近似留一法公式ALO的一大类正则化估计。给定正则化估计,计算ALO需要较小的计算开销。通过对数据生成过程的微小假设,我们得到了留一交叉验证和近似留一交叉验证之间差异的有限样本上界,|LO-ALO|.我们的理论分析表明,|LO-ALO|-> 0,当n,p ->无穷大时,其中特征向量的维数p可以与观测值的数量n相当或甚至更大。尽管问题的高维性,我们的理论结果不需要任何稀疏性假设的回归系数向量。我们大量的数值实验表明,|LO-ALO|随着nandp的增加而减小,揭示了近似留一交叉验证的优秀有限样本性能。我们进一步说明了我们提出的样本外风险估计方法的有用性的一个例子,从大鼠内侧内嗅皮层的空间敏感神经元(网格细胞)的真实的记录。
The paper considers the problem of out-of-sample risk estimation under the high dimensional settings where standard techniques such asK-fold cross-validation suffer from large biases. Motivated by the low bias of the leave-one-out cross-validation method, we propose a computationally efficient closed form approximate leave-one-out formula ALO for a large class of regularized estimators. Given the regularized estimate, calculating ALO requires a minor computational overhead. With minor assumptions about the data-generating process, we obtain a finite sample upper bound for the difference between leave-one-out cross-validation and approximate leave-one-out cross-validation, |LO-ALO|. Our theoretical analysis illustrates that |LO-ALO|-> 0 with overwhelming probability, whenn,p ->infinity, where the dimensionpof the feature vectors may be comparable with or even greater than the number of observations,n. Despite the high dimensionality of the problem, our theoretical results do not require any sparsity assumption on the vector of regression coefficients. Our extensive numerical experiments show that |LO-ALO| decreases asnandpincrease, revealing the excellent finite sample performance of approximate leave-one-out cross-validation. We further illustrate the usefulness of our proposed out-of-sample risk estimation method by an example of real recordings from spatially sensitive neurons (grid cells) in the medial entorhinal cortex of a rat.