Experience and Practice of Batch Scheduling on Leadership Supercomputers at Argonne

Experience and Practice of Batch Scheduling on Leadership Supercomputers at Argonne
复制标题

阿贡领导型超级计算机批量调度的经验与实践

DOI:
10.1007/978-3-319-77398-8_1
复制
发表时间:
2017
期刊:
2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Z. Lan
Z. Lan
中科院分区:
--
文献类型:
--
作者:
W. Allcock;Paul M. Rich;Yuping Fan;Z. Lan

文献摘要

被引文献

相似文献

美国能源部阿贡领导计算设施(ALCF)的使命是通过与计算科学界合作,设计和提供世界领先的计算设施,加速人类的重大科学发现和工程突破。ALCF运行的超级计算机通常是世界上最快的5台计算机之一。具体来说,ALCF正在寻找的科学要么太大而无法在其他地方运行,要么需要很长时间而不切实际(即,“能力工作”)。在ALCF中,批调度对于在一组约束条件下实现一组站点目标起着关键作用。虽然系统利用率是ALCF的一个重要目标,但其最大的任务限制是使极端规模的并行作业优先。在本文中,我们将描述具体的调度目标和约束,分析从48机架千兆级超级计算机Mira收集的2013-2017年工作负载轨迹,并讨论ALCF即将面临的调度挑战。
The mission of the DOE Argonne Leadership Computing Facility (ALCF) is to accelerate major scientific discoveries and engineering breakthroughs for humanity by designing and providing world-leading computing facilities in partnership with the computational science community. The ALCF operates supercomputers that are generally amongst the Top 5 fastest machines in the world. Specifically, ALCF is looking for the science that is either too big to run anywhere else, or it would take so long as to be impractical (i.e., “capability jobs”). At ALCF, batch scheduling plays a critical role for achieving a set of site goals within a set of constraints. While system utilization is an important goal at ALCF, its largest mission constraint is to enable extreme scale parallel jobs to take precedence. In this paper, we will describe the specific scheduling goals and constraints, analyze the workload traces collected in 2013–2017 from the 48-rack petascale supercomputer Mira, and discuss the upcoming scheduling challenges at ALCF.