PROV-IO: An I/O-Centric Provenance Framework for Scientific Data on HPC Systems

PROV-IO: An I/O-Centric Provenance Framework for Scientific Data on HPC Systems
复制标题

PROV-IO:HPC 系统上以 I/O 为中心的科学数据来源框架

DOI:
10.1145/3502181.3531477
复制
发表时间:
2022
期刊:
Proceedings of the 31st International Symposium on High-Performance Parallel and Distributed Computing (HPDC
影响因子:
--
通讯作者:
Zheng, Mai
Zheng, Mai
中科院分区:
--
文献类型:
--
作者:
Han, Runzhou;Byna, Suren;Tang, Houjun;Dong, Bin;Zheng, Mai

文献摘要

参考文献

被引文献

相似文献

c数据来源或数据沿袭描述了数据的生命周期。在 HPC 系统的科学工作流程中,科学家经常寻求不同的来源(例如数据产品的起源、数据集的使用模式)。 Unfortunately, existing provenance solutions cannot address the challenges due to their incompatible provenance models and/or system implementations.In this paper, we analyze three representative scientific workflows in collaboration with the domain scientists to identify concrete provenance needs.基于第一手的分析,我们提出了一个名为 PROV-IO 的起源框架,其中包括一个以 I/O 为中心的起源模型,用于精确描述科学数据以及相关的 I/O 操作和环境。此外,我们构建了 PROV-IO 原型,只需很少的手动工作即可在真实 HPC 系统上实现端到端来源支持。 PROV-IO 框架提供了选择各种来源类别的灵活性。我们对实际工作流程的实验表明,PROV-IO 可以以合理的性能有效地满足领域科学家的来源需求(例如,大多数实验的跟踪开销低于 3.5%)。此外,在我们的实验中,PROV-IO 的性能优于最先进的系统(即 ProvLake)。
cData provenance, or data lineage, describes the life cycle of data. In scientific workflows on HPC systems, scientists often seek diverse provenance (e.g., origins of data products, usage patterns of datasets). Unfortunately, existing provenance solutions cannot address the challenges due to their incompatible provenance models and/or system implementations.In this paper, we analyze three representative scientific workflows in collaboration with the domain scientists to identify concrete provenance needs. Based on the first-hand analysis, we propose a provenance framework called PROV-IO, which includes an I/O-centric provenance model for describing scientific data and the associated I/O operations and environments precisely. Moreover, we build a prototype of PROV-IO to enable end-to-end provenance support on real HPC systems with little manual effort. The PROV-IO framework provides flexibility in selecting various classes of provenance. Our experiments with realistic workflows show that PROV-IO can address the provenance needs of the domain scientists effectively with reasonable performance (e.g., less than 3.5% tracking overhead for most experiments). Moreover, PROV-IO outperforms a state-of-the-art system (i.e., ProvLake) in our experiments.
超越出处:用基于模式的平衡解释查询答案
DOI: 10.1145/3299869.3300066
发表时间: 2019
期刊: SIGMOD
影响因子: --
作者:
Miao, Zhengjie;Zeng, Qitian;Glavic, Boris;Roy, Sudeepa
通讯作者: Roy, Sudeepa
体验 ProvLake 管理 AI 工作流程的数据沿袭
DOI: --
发表时间: 2020
期刊: Anais Estendidos do XVI Simpósio Brasileiro de Sistemas de Informação (Anais Estendidos do SBSI 2020)
影响因子: --
作者:
L. Azevedo;Renan Souza;R. Thiago;Elton F. S. Soares;M. Moreno
通讯作者: M. Moreno
Lustre 文件系统检查器的性能研究:瓶颈和潜力
DOI: 10.1109/msst.2019.00-20
发表时间: 2019
期刊: 2019 35th Symposium on Mass Storage Systems and Technologies (MSST
影响因子: --
作者:
Dai, Dong;Gatla, Om Rameshwar;Zheng, Mai
通讯作者: Zheng, Mai
端到端电子科学:在海洋观测站集成工作流程、查询、可视化和来源
DOI: --
发表时间: 2008
期刊: 2008 IEEE Fourth International Conference on eScience
影响因子: --
作者:
Bill Howe;Peter Lawson;Renee Bellinger;Erik W. Anderson;E. Santos;J. Freire;C. Scheidegger;A. Baptista;Cláudio T. Silva
通讯作者: Cláudio T. Silva
Komadu:科学数据来源的捕获和可视化系统
DOI: 10.5334/jors.bq
发表时间: 2015
影响因子: --
作者:
Isuru Suriarachchi;Quan Zhou;Beth Plale
通讯作者: Beth Plale