Experiencing ProvLake to Manage the Data Lineage of AI Workflows

Experiencing ProvLake to Manage the Data Lineage of AI Workflows
复制标题

体验 ProvLake 管理 AI 工作流程的数据沿袭

DOI:
--
复制
发表时间:
2020
期刊:
Anais Estendidos do XVI Simpósio Brasileiro de Sistemas de Informação (Anais Estendidos do SBSI 2020)
影响因子:
--
通讯作者:
M. Moreno
M. Moreno
中科院分区:
--
文献类型:
--
作者:
L. Azevedo;Renan Souza;R. Thiago;Elton F. S. Soares;M. Moreno

文献摘要

被引文献

相似文献

机器学习(ML)是人工智能系统背后的核心概念,它由数据驱动并生成ML模型。这些模型用于决策制定,并且通过以下方式信任其输出至关重要,例如,理解产生它们的过程。解释ML模型的一种方法是通过跟踪整个ML生命周期,生成其数据谱系,这可以通过起源数据管理技术来完成。在这项工作中,我们提出了使用ProvLake工具在ML生命周期中进行ML源数据管理,用于井顶拾取,这是石油和天然气勘探的一个重要过程。我们展示了ProvLake如何支持ML模型的验证,了解ML模型是否根据领域特征进行泛化,以及它们的推导。
Machine Learning (ML) is a core concept behind Artificial Intelligence systems, which work driven by data and generate ML models. These models are used for decision making, and it is crucial to trust their outputs by, e.g., understanding the process that derives them. One way to explain the derivation of ML models is by tracking the whole ML lifecycle, generating its data lineage, which may be accomplished by provenance data management techniques. In this work, we present the use of ProvLake tool for ML provenance data management in the ML lifecycle for Well Top Picking, an essential process in Oil and Gas exploration. We show how ProvLake supported the validation of ML models, the understanding of whether the ML models generalize respecting the domain characteristics, and their derivation.