Probabilistic advanced reservations for batch-scheduled parallel machines

Probabilistic advanced reservations for batch-scheduled parallel machines
复制标题

批量调度并行机的概率提前预留

DOI:
--
复制
发表时间:
2008
期刊:
ACM SIGPLAN Symposium on Principles & Practice of Parallel Programming
影响因子:
--
通讯作者:
J. Brevik
J. Brevik
中科院分区:
--
文献类型:
--
作者:
Daniel Nurmi;R. Wolski;J. Brevik

文献摘要

被引文献

相似文献

在高性能计算(HPC)环境中,多处理器机器在具有潜在竞争资源需求的用户之间共享,使用空间共享将处理器分配给用户工作负载。通常,用户通过将其作业提交给集中式批处理调度器来与给定的机器交互,该中央批处理调度器实施特定于站点且通常部分隐藏的策略,该策略旨在最大化机器利用率,同时提供可容忍的周转时间。在实践中,虽然大多数HPC系统具有良好的利用率水平,但单个作业等待开始执行的时间量已被证明是高度可变的,难以预测,从而导致用户困惑和/或沮丧。已提出的处理这种不确定性的一种方法是允许愿意提前计划的用户对处理器资源进行“提前预订”。然而,到目前为止,很少有HPC中心为其一般用户群提供高级预订功能,因为他们担心(得到之前的研究支持),如果引入高级预订,机器利用率将会降低。在这项工作中,我们描述了Varq,一种新的作业调度方法,它只使用现有的尽力而为批处理调度器和策略为用户提供概率“虚拟”的提前预留。Varq充当覆盖层,提交与调度器服务的正常工作负载没有区别的作业。我们描述了我们用来实施VARQ的统计方法,详细介绍了它在一些高性能计算环境中的有效性的经验评估,并探索了如果VARQ被广泛使用的潜在的未来影响。在不要求HPC站点支持高级预留的情况下,我们发现Varq可以概率地实现预留能力,并且这种概率方法的效果不太可能对资源利用率产生负面影响。
In high-performance computing (HPC) settings, in which multiprocessor machines are shared among users with potentially competing resource demands, processors are allocated to user workload using space sharing. Typically, users interact with a given machine by submitting their jobs to a centralized batch scheduler that implements a site-specific, and often partially hidden, policy designed to maximize machine utilization while providing tolerable turn-around times. In practice, while most HPC systems experience good utilization levels, the amount of time experienced by individual jobs waiting to begin execution has been shown to be highly variable and difficult to predict, leading to user confusion and/or frustration. One method for dealing with this uncertainty that has been proposed is to allow users who are willing to plan ahead to make “advanced reservations” for processor resources. To date, however, few if any HPC centers provide an advanced reservation capability to their general user populations for fear (supported by previous research) that diminished machine utilization will occur if and when advanced reservations are introduced. In this work, we describe VARQ, a new method for job scheduling that provides users with probabilistic “virtual” advanced reservations using only existing best effort batch schedulers and policies. VARQ functions as an overlay, submitting jobs that are indistinguishable from the normal workload serviced by a scheduler. We describe the statistical methods we use to implement VARQ, detail an empirical evaluation of its effectiveness in a number of HPC settings, and explore the potential future impact of VARQ should it become widely used. Without requiring HPC sites to support advanced reservations, we find that VARQ can implement a reservation capability probabilistically and that the effects of this probabilistic approach are unlikely to negatively affect resource utilization.