Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine Learning

Beta Shapley: a Unified and Noise-reduced Data Valuation Framework for Machine Learning
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
Yongchan Kwon;James Y. Zou
Yongchan Kwon;James Y. Zou
中科院分区:
其他
文献类型:
--
作者:
Yongchan Kwon;James Y. Zou

文献摘要

相似文献

数据Shapley最近被提出作为一个原则性的框架,以量化个体数据在机器学习中的贡献。它可以有效地识别对学习算法有用或有害的数据点。在本文中,我们提出了Beta Shapley,这是一个实质性的推广数据Shapley。Beta Shapley通过放松Shapley值的效率公理而自然产生,这对于机器学习设置并不重要。Beta Shapley统一了几种流行的数据估值方法,并将数据Shapley作为特例。此外,我们证明Beta Shapley具有几个理想的统计特性,并提出了有效的算法来估计它。我们证明Beta Shapley在几个下游ML任务上优于最先进的数据估值方法,例如:1)检测错误标记的训练数据; 2)使用子样本学习; 3)识别添加或删除对模型产生最大积极或消极影响的点。
Data Shapley has recently been proposed as a principled framework to quantify the contribution of individual datum in machine learning. It can effectively identify helpful or harmful data points for a learning algorithm. In this paper, we propose Beta Shapley, which is a substantial generalization of Data Shapley. Beta Shapley arises naturally by relaxing the efficiency axiom of the Shapley value, which is not critical for machine learning settings. Beta Shapley unifies several popular data valuation methods and includes data Shapley as a special case. Moreover, we prove that Beta Shapley has several desirable statistical properties and propose efficient algorithms to estimate it. We demonstrate that Beta Shapley outperforms state-of-the-art data valuation methods on several downstream ML tasks such as: 1) detecting mislabeled training data; 2) learning with subsamples; and 3) identifying points whose addition or removal have the largest positive or negative impact on the model.