Flux: A Next-Generation Resource Management Framework for Large HPC Centers

Flux: A Next-Generation Resource Management Framework for Large HPC Centers
复制标题

Flux:大型 HPC 中心的下一代资源管理框架

DOI:
10.1109/icppw.2014.15
复制
发表时间:
2014
期刊:
2014 43rd International Conference on Parallel Processing Workshops
影响因子:
--
通讯作者:
M. Schulz
M. Schulz
中科院分区:
--
文献类型:
--
作者:
D. Ahn;J. Garlick;Mark Grondona;D. Lipari;B. Springmeyer;M. Schulz

文献摘要

被引文献

相似文献

资源和作业管理软件对于高效执行应用程序的高性能计算(HPC)至关重要。然而,由于不断增加的系统规模、资源和工作负载多样性、各种资源之间的相互作用(例如,在计算集群和全局文件系统之间),以及资源约束的复杂性,诸如严格的功率预算。为了解决这一差距,我们提出了通量,可扩展的工作和资源管理框架,专门设计用于处理下一代HPC中心的要求。Flux将整个计算设施作为不同资源集的一个公共池,使得设施能够适应站点范围的约束(例如,功率限制)。然而,它的可扩展和分布式设计仍然提供了可扩展和有效的调度策略。本文详细介绍了Flux的设计,并描述和评估了我们对关键运行时组件的初始原型设计工作。我们的结果表明,我们的运行时原型提供了强大的和可预测的可扩展性.
Resource and job management software is crucial to High Performance Computing (HPC) for efficient application execution. However, current systems and approaches can no longer keep up with the challenges large HPC centers are facing due to ever-increasing system scales, resource and workload diversity, interplays between various resources (e.g., between compute clusters and a global file system), and complexity of resource constraints such as strict power budgeting. To address this gap, we propose Flux, an extensible job and resource management framework specifically designed to deal with the requirements of next-generation HPC centers. Flux targets an entire computing facility as one common pool of diverse sets of resources, enabling the facility to accommodate site-wide constraints (e.g., for power limits). Yet, its scalable and distributed design still offers scalable and effective scheduling strategies. This paper details the design of Flux and describes and evaluates our initial prototyping effort of the key run-time components. Our results show that our run- time prototype provides strong and predictable scalability.