Lazy Estimation of Variable Importance for Large Neural Networks

Lazy Estimation of Variable Importance for Large Neural Networks
复制标题

DOI:
10.48550/arxiv.2207.09097
复制
发表时间:
2022-07
期刊:
--
影响因子:
--
通讯作者:
Yue Gao;Abby Stevens;Garvesh Raskutti;R. Willett
Yue Gao;Abby Stevens;Garvesh Raskutti;R. Willett
中科院分区:
其他
文献类型:
--
作者:
Yue Gao;Abby Stevens;Garvesh Raskutti;R. Willett

文献摘要

被引文献

相似文献

随着不透明的预测模型越来越多地影响现代生活的许多领域,人们对量化给定输入变量对做出特定预测的重要性的兴趣越来越大。最近,测量变量重要性(VI)的模型不可知方法激增,这些方法分析了在所有变量上训练的完整模型与排除感兴趣变量的简化模型之间预测能力的差异。这些方法的一个共同瓶颈是对每个变量(或变量子集)的简化模型进行估计,这是一个昂贵的过程,通常没有理论上的保证。在这项工作中,我们提出了一种快速灵活的方法来逼近具有重要推理保证的约简模型。我们用在全模型参数上初始化的线性化取代了对广义神经网络进行完全再训练的需要。通过添加脊状惩罚使问题凸出,我们证明当脊状惩罚参数足够大时,我们的方法估计变量重要性度量的错误率为$O(\frac{1}{\sqrt{n}})$,其中$n$为训练样本的数量。我们还证明了我们的估计量是渐近正态的,使我们能够为VI估计提供置信界限。我们通过模拟证明了我们的方法在几种数据生成机制下是快速和准确的,我们在一个季节性气候预测的例子上证明了它在现实世界中的适用性。
As opaque predictive models increasingly impact many areas of modern life, interest in quantifying the importance of a given input variable for making a specific prediction has grown. Recently, there has been a proliferation of model-agnostic methods to measure variable importance (VI) that analyze the difference in predictive power between a full model trained on all variables and a reduced model that excludes the variable(s) of interest. A bottleneck common to these methods is the estimation of the reduced model for each variable (or subset of variables), which is an expensive process that often does not come with theoretical guarantees. In this work, we propose a fast and flexible method for approximating the reduced model with important inferential guarantees. We replace the need for fully retraining a wide neural network by a linearization initialized at the full model parameters. By adding a ridge-like penalty to make the problem convex, we prove that when the ridge penalty parameter is sufficiently large, our method estimates the variable importance measure with an error rate of $O(\frac{1}{\sqrt{n}})$ where $n$ is the number of training samples. We also show that our estimator is asymptotically normal, enabling us to provide confidence bounds for the VI estimates. We demonstrate through simulations that our method is fast and accurate under several data-generating regimes, and we demonstrate its real-world applicability on a seasonal climate forecasting example.