Modeling and Analyzing Waiting Policies for Cloud-Enabled Schedulers

Modeling and Analyzing Waiting Policies for Cloud-Enabled Schedulers
复制标题

DOI:
10.1109/tpds.2021.3086270
复制
发表时间:
2021-12
影响因子:
5.3
通讯作者:
Pradeep Ambati;Noman Bashir;David E. Irwin;Prashant J. Shenoy
Pradeep Ambati;Noman Bashir;David E. Irwin;Prashant J. Shenoy
中科院分区:
计算机科学2区
文献类型:
--
作者:
Pradeep Ambati;Noman Bashir;David E. Irwin;Prashant J. Shenoy

文献摘要

相似文献

云平台已经普及了云结构即服务(IaaS)购买模型,该模型使用户能够按需租用计算资源来执行其作业。但是,如果资源利用率高,购买固定资源仍然比租用便宜得多。因此,为了优化成本,用户必须根据他们的工作负载来决定提供多少固定资源,而不是“按需”租用多少固定资源。在这篇文章中,我们介绍了等待策略的概念,并表明最佳成本取决于它。等待策略显式地控制作业等待资源的时间,因为作业永远不需要等待,因为云平台提供了无限可扩展性的假象。等待策略是调度策略的对偶:调度策略决定当固定资源可用时应该运行哪些作业,而等待策略决定当固定资源不可用时应该等待哪些作业。我们定义了多种等待策略,并开发了简单而通用的分析模型,以揭示它们在固定资源配置、成本和作业等待时间之间的权衡。我们评估了不同的等待策略对运行在14. 3k核集群上的由14M作业组成的真实的长达一年的批处理工作负载的影响。我们发现,一个复合等待政策,这迫使长运行时间或短等待时间的作业等待固定的资源,提供了最佳的权衡。与当前集群相比,该策略降低了成本(5%)和平均作业等待时间(7倍),与租用按需资源相比,还降低了成本(43%),平均作业等待时间略有增加(1.74小时)。
Cloud platforms have popularized the Infrastructure-as-a-Service (IaaS) purchasing model, which enables users to rent computing resources on demand to execute their jobs. However, buying fixed resources is still much cheaper than renting if their resource utilization is high. Thus, to optimize cost, users must decide how many fixed resources to provision versus rent “on demand” based on their workload. In this article, we introduce the concept of a waiting policy for cloud-enabled schedulers and show that the optimal cost depends on it. The waiting policy explicitly controls how long jobs wait for resources, as jobs never need to wait, since cloud platforms provide the illusion of infinite scalability. A waiting policy is the dual of a scheduling policy: while a scheduling policy determines which jobs should run when fixed resources are available, a waiting policy determines which jobs should wait when fixed resources are not available. We define multiple waiting policies and develop simple and general analytical models to reveal their tradeoff between fixed resource provisioning, cost, and job waiting time. We evaluate the impact of different waiting policies on a real year-long batch workload consisting of 14M jobs run on a 14.3k-core cluster. We show that a compound waiting policy, which forces jobs with long running times or short waiting times to wait for fixed resources, offers the best tradeoff. The policy decreases both the cost (by 5 percent) and mean job waiting time (by 7×) compared to the current cluster, and also decreases the cost (by 43 percent) compared to renting on-demand resources for a modest increase in mean job waiting time (at 1.74 hours).