Modeling The Temporally Constrained Preemptions of Transient Cloud VMs

Modeling The Temporally Constrained Preemptions of Transient Cloud VMs
复制标题

DOI:
10.1145/3369583.3392671
复制
发表时间:
2020-06
期刊:
Proceedings of the 29th International Symposium on High-Performance Parallel and Distributed Computing
影响因子:
--
通讯作者:
Jcs Kadupitige;V. Jadhao;Prateek Sharma
Jcs Kadupitige;V. Jadhao;Prateek Sharma
中科院分区:
其他
文献类型:
--
作者:
Jcs Kadupitige;V. Jadhao;Prateek Sharma

文献摘要

相似文献

Amazon Spot实例、Google Preemptible VM和Azure低优先级批处理VM等瞬态云服务器可以将云计算成本降低多达10倍,但可以被云提供商单方面抢占。了解抢占特性(如频率)是最小化抢占对应用程序性能、可用性和成本的影响的关键第一步。然而,很少有人了解时间约束的抢占-其中抢占必须发生在一个给定的时间窗口。我们通过对Google的可抢占虚拟机(最大生命周期为24小时)进行大规模实证研究来研究时间受限的抢占,开发了一个新的抢占概率模型,新的模型驱动的资源管理策略,并在科学计算工作负载的批处理计算服务中实现它们。我们的统计和实验分析表明,时间约束的抢占不是均匀分布的,而是依赖于时间并且具有浴缸形状。我们发现,现有的无记忆模型和政策是不适合的时间约束抢占。我们建立了一个新的浴缸抢占的概率模型,并通过可靠性理论的透镜进行分析。为了突出我们的模型的有效性,我们制定了优化的政策,作业调度和检查点。与现有技术相比,我们基于模型的策略可以将作业失败的概率降低2倍以上。我们还将我们的策略作为科学计算应用程序的批处理计算服务的一部分来实施,与传统云部署相比,这将成本降低了5倍,并将性能开销保持在3%以下。
Transient cloud servers such as Amazon Spot instances, Google Preemptible VMs, and Azure Low-priority batch VMs, can reduce cloud computing costs by as much as 10x, but can be unilaterally preempted by the cloud provider. Understanding preemption characteristics (such as frequency) is a key first step in minimizing the effect of preemptions on application performance, availability, and cost. However, little is understood about temporally constrained preemptions---wherein preemptions must occur in a given time window. We study temporally constrained preemptions by conducting a large scale empirical study of Google's Preemptible VMs (that have a maximum lifetime of 24 hours), develop a new preemption probability model, new model-driven resource management policies, and implement them in a batch computing service for scientific computing workloads. Our statistical and experimental analysis indicates that temporally constrained preemptions are not uniformly distributed but are time-dependent and have a bathtub shape. We find that existing memoryless models and policies are not suitable for temporally constrained preemptions. We develop a new probability model for bathtub preemptions and analyze it through the lens of reliability theory. To highlight the effectiveness of our model, we develop optimized policies for job scheduling and checkpointing. Compared to existing techniques, our model-based policies can reduce the probability of job failure by more than 2x. We also implement our policies as part of a batch computing service for scientific computing applications, which reduces cost by 5x compared to conventional cloud deployments and keeps performance overheads under 3%.