LAVA: Data Valuation without Pre-Specified Learning Algorithms

LAVA: Data Valuation without Pre-Specified Learning Algorithms
复制标题

DOI:
10.48550/arxiv.2305.00054
复制
发表时间:
2023-04
期刊:
ArXiv
影响因子:
--
通讯作者:
H. Just;Feiyang Kang;Jiachen T. Wang;Yi Zeng;Myeongseob Ko;Ming Jin;R. Jia
H. Just;Feiyang Kang;Jiachen T. Wang;Yi Zeng;Myeongseob Ko;Ming Jin;R. Jia
中科院分区:
其他
文献类型:
--
作者:
H. Just;Feiyang Kang;Jiachen T. Wang;Yi Zeng;Myeongseob Ko;Ming Jin;R. Jia

文献摘要

被引文献

相似文献

传统上,数据估值(DV)是作为在培训数据之间公平地将学习算法的验证性能分解的问题。结果,计算出的数据值取决于基础学习算法的许多设计选择。但是,对于许多DV用例,这种依赖性是不可取的,例如在数据采集过程中将优先级确定优先级,并在数据市场中告知定价机制。在这些情况下,需要在实际分析之前对数据进行评估,并且仍然不确定学习算法的选择。依赖性的另一个副作用是,要评估单个点的价值,需要有和没有点的学习算法来重新运行学习算法,这会造成巨大的计算负担。这项工作通过引入一个新框架,可以以一种忽略下游学习算法来重视训练数据的新框架,超越了数据评估方法的当前限制。我们的主要结果如下。 (1)我们为基于训练和验证集之间的非规定的瓦斯汀距离而建立了与训练集相关的验证性能的代理。我们表明,在某些Lipschitz条件下,任何给定模型的验证性能的上限都表征了距离。 (2)我们开发了一种新的方法来基于阶级瓦斯汀距离的灵敏度分析来重视单个数据。重要的是,这些值可以直接从计算距离时从现成优化求解器的输出中免费获得。 (3)我们评估了与检测低质量数据相关的各种用例中的新数据评估框架,并表明,令人惊讶的是,我们框架的学习 - 不合Snostic特征使SOTA性能可以显着改善,同时更快地订单。
Traditionally, data valuation (DV) is posed as a problem of equitably splitting the validation performance of a learning algorithm among the training data. As a result, the calculated data values depend on many design choices of the underlying learning algorithm. However, this dependence is undesirable for many DV use cases, such as setting priorities over different data sources in a data acquisition process and informing pricing mechanisms in a data marketplace. In these scenarios, data needs to be valued before the actual analysis and the choice of the learning algorithm is still undetermined then. Another side-effect of the dependence is that to assess the value of individual points, one needs to re-run the learning algorithm with and without a point, which incurs a large computation burden. This work leapfrogs over the current limits of data valuation methods by introducing a new framework that can value training data in a way that is oblivious to the downstream learning algorithm. Our main results are as follows. (1) We develop a proxy for the validation performance associated with a training set based on a non-conventional class-wise Wasserstein distance between training and validation sets. We show that the distance characterizes the upper bound of the validation performance for any given model under certain Lipschitz conditions. (2) We develop a novel method to value individual data based on the sensitivity analysis of the class-wise Wasserstein distance. Importantly, these values can be directly obtained for free from the output of off-the-shelf optimization solvers when computing the distance. (3) We evaluate our new data valuation framework over various use cases related to detecting low-quality data and show that, surprisingly, the learning-agnostic feature of our framework enables a significant improvement over SOTA performance while being orders of magnitude faster.