Enhancing and abstracting scientific workflow provenance for data publishing

Enhancing and abstracting scientific workflow provenance for data publishing
复制标题

DOI:
10.1145/2457317.2457370
复制
发表时间:
2013-03
期刊:
2009 3rd International Conference on Bioinformatics and Biomedical Engineering
影响因子:
--
通讯作者:
Pinar Alper;Khalid Belhajjame;C. Goble;Pinar Senkul
Pinar Alper;Khalid Belhajjame;C. Goble;Pinar Senkul
中科院分区:
其他
文献类型:
--
作者:
Pinar Alper;Khalid Belhajjame;C. Goble;Pinar Senkul

文献摘要

被引文献

相似文献

许多科学家正在使用工作流程来系统地设计和运行计算实验。一旦工作流程被执行,科学家可能想要发布作为结果生成的数据集,以供其他科学家重复使用作为他们实验的输入。在此过程中,科学家需要通过指定描述数据集的元数据信息来管理此类数据集,例如它的衍生历史、起源和所有权。为了帮助科学家完成这项任务,我们在本文中探讨了在制定工作流时工作流管理系统收集的来源跟踪的使用。具体来说,我们确定了这种原始来源跟踪在支持数据发布任务方面的缺点,并提出了一种方法,通过该方法可以导出适合数据发布任务的精炼但信息更丰富的来源跟踪。
Many scientists are using workflows to systematically design and run computational experiments. Once the workflow is executed, the scientist may want to publish the dataset generated as a result, to be, e.g., reused by other scientists as input to their experiments. In doing so, the scientist needs to curate such dataset by specifying metadata information that describes it, e.g. its derivation history, origins and ownership. To assist the scientist in this task, we explore in this paper the use of provenance traces collected by workflow management systems when enacting workflows. Specifically, we identify the shortcomings of such raw provenance traces in supporting the data publishing task, and propose an approach whereby distilled, yet more informative, provenance traces that are fit for the data publishing task can be derived.