Efficient computation and analysis of distributional Shapley values

Efficient computation and analysis of distributional Shapley values
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
--
影响因子:
--
通讯作者:
Yongchan Kwon;Manuel A. Rivas;James Y. Zou
Yongchan Kwon;Manuel A. Rivas;James Y. Zou
中科院分区:
其他
文献类型:
--
作者:
Yongchan Kwon;Manuel A. Rivas;James Y. Zou

文献摘要

相似文献

最近,分布数据Shapley值(DShapley Value,DShapley Value)被提出作为一个原则性框架来量化机器学习中个体数据的贡献。DShapley将Shapley值的基本博弈论概念发展成一个统计框架,并可用于识别对学习算法有用(或有害)的数据点。然而,估计DShapley的计算代价很高,这可能是在实践中使用它的主要挑战。此外,很少有关于这个值如何取决于数据特征的数学分析。本文给出了线性回归和非参数密度估计的正则问题的DShapley的第一个解析表达式。这些解析形式提供了计算DShapley的新算法,比以前最先进的算法快了几个数量级。此外,我们的公式是可直接解释的,并提供了对不同类型数据的值如何变化的量化见解。我们展示了我们的DShapley方法在多个真实和合成数据集上的有效性。
Distributional data Shapley value (DShapley) has been recently proposed as a principled framework to quantify the contribution of individual datum in machine learning. DShapley develops the foundational game theory concept of Shapley values into a statistical framework and can be applied to identify data points that are useful (or harmful) to a learning algorithm. Estimating DShapley is computationally expensive, however, and this can be a major challenge to using it in practice. Moreover, there has been little mathematical analyses of how this value depends on data characteristics. In this paper, we derive the first analytic expressions for DShapley for the canonical problems of linear regression and non-parametric density estimation. These analytic forms provide new algorithms to compute DShapley that are several orders of magnitude faster than previous state-of-the-art. Furthermore, our formulas are directly interpretable and provide quantitative insights into how the value varies for different types of data. We demonstrate the efficacy of our DShapley approach on multiple real and synthetic datasets.