Applying the Virtual Data Provenance Model

Applying the Virtual Data Provenance Model
复制标题

DOI:
10.1007/11890850_16
复制
发表时间:
2006-05
期刊:
--
影响因子:
--
通讯作者:
Yong Zhao;M. Wilde;Ian T Foster
Yong Zhao;M. Wilde;Ian T Foster
中科院分区:
其他
文献类型:
--
作者:
Yong Zhao;M. Wilde;Ian T Foster

文献摘要

被引文献

相似文献

在科学、工程和商业的许多领域,数据分析系统被用于从描述实验结果或模拟现象的数据集中获得新数据(最终,人们希望是知识)。为了支持这种分析,我们开发了一个“虚拟数据系统”,它允许用户首先定义,然后调用,最后探索执行这种数据派生的过程(以及包含多个过程调用的工作流)的来源。底层执行模型是“功能性的”,因为过程读取(但不修改)它们的输入,并通过确定性计算产生输出。这个属性使得虚拟数据系统不仅可以直接记录生成任何给定数据对象的配方,还可以记录关于执行配方的环境的足够信息,所有这些都具有足够的保真度,以便可以重新执行用于创建数据对象的步骤,以便在以后的时间或不同的位置重新生成数据对象。虚拟数据系统将这些信息与语义注释一起保存在一个集成的模式中,从而支持强大的查询功能,其中可以利用数据派生过程结构知识所隐含的丰富语义信息来提供融合配方、历史和特定于应用程序的语义的信息环境。我们在这里概述了这种集成、它支持的查询和转换,以及这些功能如何为科学流程服务的示例。
In many domains of science, engineering, and commerce, data analysis systems are employed to derive new data (and ultimately, one hopes, knowledge) from datasets describing experimental results or simulated phenomena. To support such analyses, we have developed a “virtual data system” that allows users first to define, then to invoke, and finally explore the provenance of procedures (and workflows comprising multiple procedure calls) that perform such data derivations. The underlying execution model is “functional” in the sense that procedures read (but do not modify) their input and produce output via deterministic computations. This property makes it straightforward for the virtual data system to record not only the recipe for producing any given data object but also sufficient information about the environment in which the recipe has been executed, all with sufficient fidelity that the steps used to create a data object can be re-executed to reproduce the data object at a later time or a different location. The virtual data system maintains this information in an integrated schema alongside semantic annotations, and thus enables a powerful query capability in which the rich semantic information implied by knowledge of the structure of data derivation procedures can be exploited to provide an information environment that fuses recipe, history, and application-specific semantics. We provide here an overview of this integration, the queries and transformations that it enables, and examples of how these capabilities can serve scientific processes.