Capturing and querying fine-grained provenance of preprocessing pipelines in data science

Capturing and querying fine-grained provenance of preprocessing pipelines in data science
复制标题

DOI:
10.14778/3436905.3436911
复制
发表时间:
2020-12
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Adriane P. Chapman;P. Missier;G. Simonelli;Riccardo Torlone
Adriane P. Chapman;P. Missier;G. Simonelli;Riccardo Torlone
中科院分区:
其他
文献类型:
--
作者:
Adriane P. Chapman;P. Missier;G. Simonelli;Riccardo Torlone

文献摘要

被引文献

相似文献

设计用于清理、转换和更改数据以准备学习预测模型的数据处理管道会影响这些模型的准确性和性能,以及其他属性,例如模型公平性。因此,重要的是为开发人员提供深入了解管道步骤如何影响数据的方法,从原始输入到准备用于学习的训练集。虽然其他工作会跟踪关系运算符管道的创建和更改,但在这项工作中,我们分析了机器学习过程中数据准备的典型操作,并提供基础设施,用于在数据集内的各个元素级别生成非常细粒度的出处记录。我们的贡献包括:(i)预处理操作符的核心集的正式定义,以及每个操作符的起源模式的定义,以及(ii)与Python一起工作的应用程序级起源捕获库的原型实现。我们报告了在真实的ML基准测试管道和TCP-DI上进行的起源处理和存储开销以及可扩展性实验,并展示了如何使用产生的起源来回答一套起源基准测试查询,这些查询支持开发人员的一些调试问题,如Data Science Stack Exchange上所表达的。
Data processing pipelines that are designed to clean, transform and alter data in preparation for learning predictive models, have an impact on those models' accuracy and performance, as well on other properties, such as model fairness. It is therefore important to provide developers with the means to gain an in-depth understanding of how the pipeline steps affect the data, from the raw input to training sets ready to be used for learning. While other efforts track creation and changes of pipelines of relational operators, in this work we analyze the typical operations of data preparation within a machine learning process, and provide infrastructure for generating very granular provenance records from it, at the level of individual elements within a dataset. Our contributions include: (i) the formal definition of a core set of preprocessing operators, and the definition of provenance patterns for each of them, and (ii) a prototype implementation of an application-level provenance capture library that works alongside Python. We report on provenance processing and storage overhead and scalability experiments, carried out over both real ML benchmark pipelines and over TCP-DI, and show how the resulting provenance can be used to answer a suite of provenance benchmark queries that underpin some of the developers' debugging questions, as expressed on the Data Science Stack Exchange.