Enabling large-scale scientific workflows on petascale resources using MPI master/worker

Enabling large-scale scientific workflows on petascale resources using MPI master/worker
复制标题

使用 MPI master/worker 在千万亿级资源上启用大规模科学工作流程

DOI:
10.1145/2335755.2335846
复制
发表时间:
2012
期刊:
Sci. Program.
影响因子:
--
通讯作者:
P. Maechling
P. Maechling
中科院分区:
--
文献类型:
--
作者:
M. Rynge;S. Callaghan;E. Deelman;G. Juve;Gaurang Mehta;K. Vahi;P. Maechling

文献摘要

被引文献

相似文献

计算科学家通常需要执行大型,松散耦合的并行应用,例如工作流和任务袋,以进行研究。这些应用程序通常由许多短运行的序列任务组成,这些任务经常需要大量的计算和存储。为了在合理的时间内产生结果,科学家想使用Petascale资源执行这些应用程序。过去,这是一个挑战,因为Petascale系统并非旨在有效地执行此类工作量。在本文中,我们描述了一种在分布式Petascale系统上执行大型,细粒度的工作流程的新方法。我们的解决方案涉及将工作流程分配到独立的子图中,然后将每个子图作为独立的MPI作业提交给可用资源(通常是远程)。我们描述了如何在Pegasus工作流管理系统中实施分区和工作管理。我们还解释了这种方法如何为与系统体系结构,队列政策和优先级以及应用程序重用和开发相关的挑战提供端到端解决方案。最后,我们描述了如何使用该系统来实现XSEDE资源上非常大的地震危险分析应用程序。
Computational scientists often need to execute large, loosely-coupled parallel applications such as workflows and bags of tasks in order to do their research. These applications are typically composed of many, short-running, serial tasks, which frequently demand large amounts of computation and storage. In order to produce results in a reasonable amount of time, scientists would like to execute these applications using petascale resources. In the past this has been a challenge because petascale systems are not designed to execute such workloads efficiently. In this paper we describe a new approach to executing large, fine-grained workflows on distributed petascale systems. Our solution involves partitioning the workflow into independent subgraphs, and then submitting each subgraph as a self-contained MPI job to the available resources (often remote). We describe how the partitioning and job management has been implemented in the Pegasus Workflow Management System. We also explain how this approach provides an end-to-end solution for challenges related to system architecture, queue policies and priorities, and application reuse and development. Finally, we describe how the system is being used to enable the execution of a very large seismic hazard analysis application on XSEDE resources.