Efficient nonparametric statistical inference on population feature importance using Shapley values

Efficient nonparametric statistical inference on population feature importance using Shapley values
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
Proceedings of machine learning research
影响因子:
--
通讯作者:
B. Williamson;Jean Feng
B. Williamson;Jean Feng
中科院分区:
其他
文献类型:
--
作者:
B. Williamson;Jean Feng

文献摘要

相似文献

预测任务中变量的真实群体水平重要性提供了有关潜在数据生成机制的有用知识,并有助于决定在后续实验中收集哪些测量值。对这种重要性的有效统计推断是理解感兴趣人群的关键组成部分。我们提出了一个计算效率的估计和获得有效的统计推断的Shapley人口变量重要性测度(SPVIM)的程序。虽然真正的SPVIM的计算复杂度与变量的数量呈指数级,但我们提出了一种基于随机抽样的估计器,仅对给定n个观测值的Θ(n)个特征子集进行抽样。我们证明了我们的估计收敛于一个渐近最优的速度。此外,通过推导我们的估计量的渐近分布,我们构造有效的置信区间和假设检验。我们的程序在模拟中具有良好的有限样本性能,并且当应用不同的机器学习算法时,对于住院死亡率预测任务,可以产生类似的变量重要性估计。
The true population-level importance of a variable in a prediction task provides useful knowledge about the underlying data-generating mechanism and can help in deciding which measurements to collect in subsequent experiments. Valid statistical inference on this importance is a key component in understanding the population of interest. We present a computationally efficient procedure for estimating and obtaining valid statistical inference on the Shapley Population Variable Importance Measure (SPVIM). Although the computational complexity of the true SPVIM scales exponentially with the number of variables, we propose an estimator based on randomly sampling only Θ(n) feature subsets given n observations. We prove that our estimator converges at an asymptotically optimal rate. Moreover, by deriving the asymptotic distribution of our estimator, we construct valid confidence intervals and hypothesis tests. Our procedure has good finite-sample performance in simulations, and for an in-hospital mortality prediction task produces similar variable importance estimates when different machine learning algorithms are applied.