Tigres Workflow Library: Supporting Scientific Pipelines on HPC Systems

Tigres Workflow Library: Supporting Scientific Pipelines on HPC Systems
复制标题

Tigres 工作流库:支持 HPC 系统上的科学管道

DOI:
10.1109/ccgrid.2016.54
复制
发表时间:
2016
期刊:
2016 16th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing (CCGrid)
影响因子:
--
通讯作者:
L. Ramakrishnan
L. Ramakrishnan
中科院分区:
--
文献类型:
--
作者:
V. Hendrix;James Fox;D. Ghoshal;L. Ramakrishnan

文献摘要

被引文献

相似文献

科学数据量的增长导致了对新工具的需求,这些工具使用户能够操作和分析大规模资源上的数据。在过去的十年中,出现了许多科学的工作流工具。这些工具通常以分布式环境为目标,通常需要专家帮助来组成和执行工作流。数据密集型工作流通常是临时的,它们涉及迭代开发过程,包括用户在桌面上组成和测试其工作流,并向上扩展到更大的系统。本文介绍了一个支持数据密集型工作流迭代开发周期的工作流库TIGRES的设计与实现。TIGRES为一组编程模板(即顺序、并行、拆分、合并)提供应用编程接口,可用于组成和执行计算和数据流水线。我们讨论了我们对科学和合成工作流的评估结果,显示TIGRES以最小的模板开销执行(所有实验的平均值为13秒)。我们还讨论了影响高性能计算系统上科学工作流性能的各种因素(如I/O性能、执行机制)。
The growth in scientific data volumes has resulted in the need for new tools that enable users to operate on and analyze data on large-scale resources. In the last decade, a number of scientific workflow tools have emerged. These tools often target distributed environments, and often need expert help to compose and execute the workflows. Data-intensive workflows are often ad-hoc, they involve an iterative development process that includes users composing and testing their workflows on desktops, and scaling up to larger systems. In this paper, we present the design and implementation of Tigres, a workflow library that supports the iterative workflow development cycle of data-intensive workflows. Tigres provides an application programming interface to a set of programming templates i.e., sequence, parallel, split, merge, that can be used to compose and execute computational and data pipelines. We discuss the results of our evaluation of scientific and synthetic workflows showing Tigres performs with minimal template overheads (mean of 13 seconds over all experiments). We also discuss various factors (e.g., I/O performance, execution mechansims) that affect the performance of scientific workflows on HPC systems.