A Survey on Collecting, Managing, and Analyzing Provenance from Scripts

A Survey on Collecting, Managing, and Analyzing Provenance from Scripts
复制标题

DOI:
10.1145/3311955
复制
发表时间:
2019-07-01
影响因子:
16.6
通讯作者:
Braganholo, Vanessa
Braganholo, Vanessa
中科院分区:
计算机科学1区
文献类型:
--
作者:
Pimentel, Joao Felipe;Freire, Juliana;Braganholo, Vanessa

文献摘要

被引文献

相似文献

脚本被广泛用于设计和运行科学实验。脚本语言易于学习和使用,并且与传统编程语言相比,它们允许以更少的步骤指定和执行复杂的任务。然而,它们在再现性和数据管理方面也有重要的限制。随着实验的迭代改进,对每个实验运行(或试验)进行推理,跟踪试验和实验实例之间的关联以及试验之间的差异,并将结果与特定的输入数据和参数联系起来是具有挑战性的。已经提出了通过收集、管理和分析脚本的来源来解决这些限制的方法。在这篇文章中,我们调查了脚本来源的艺术状态。我们已经确定了方法,通过一个详尽的协议,向前和向后的文献滚雪球。在详细研究的基础上,提出了一种分类方法,并对该方法进行了分类。
Scripts are widely used to design and run scientific experiments. Scripting languages are easy to learn and use, and they allow complex tasks to be specified and executed in fewer steps than with traditional programming languages. However, they also have important limitations for reproducibility and data management. As experiments are iteratively refined, it is challenging to reason about each experiment run (or trial), to keep track of the association between trials and experiment instances as well as the differences across trials, and to connect results to specific input data and parameters. Approaches have been proposed that address these limitations by collecting, managing, and analyzing the provenance of scripts. In this article, we survey the state of the art in provenance for scripts. We have identified the approaches by following an exhaustive protocol of forward and backward literature snowballing. Based on a detailed study, we propose a taxonomy and classify the approaches using this taxonomy.