Optimizing the Cost of Executing Mixed Interactive and Batch Workloads on Transient VMs

Optimizing the Cost of Executing Mixed Interactive and Batch Workloads on Transient VMs
复制标题

DOI:
10.1145/3309697.3331489
复制
发表时间:
2019-06
期刊:
Abstracts of the 2019 SIGMETRICS/Performance Joint International Conference on Measurement and Modeling of Computer Systems
影响因子:
--
通讯作者:
Pradeep Ambati;David E. Irwin
Pradeep Ambati;David E. Irwin
中科院分区:
其他
文献类型:
--
作者:
Pradeep Ambati;David E. Irwin

文献摘要

被引文献

相似文献

容器编排平台(COPs),例如Kubernetes,越来越多地通过自动化容器中封装的应用程序之间的资源分配来管理大规模集群。越来越多的情况是,COPs所依赖的资源是从云平台动态获取的虚拟机(VMs)。COPs可以从云平台提供的许多不同类型的虚拟机中进行选择,这些虚拟机在成本、性能和可用性方面各不相同。虽然临时虚拟机的成本明显低于按需虚拟机,但平台可能随时撤销它们,导致它们不可用。虽然临时虚拟机的价格很有吸引力,但它们的不可靠性对于旨在支持混合工作负载的COPs来说是一个问题,这些混合工作负载不仅包括容忍延迟的批处理作业,还包括对高可用性有要求的长期运行的交互式服务。为了解决这个问题,我们设计了TR - Kubernetes,这是一种容器编排平台,它使用临时虚拟机优化在云平台上执行混合交互式和批处理工作负载的成本。为此,TR - Kubernetes通过在大多数情况下获取比实际所需多得多的临时虚拟机,来满足交互式服务指定的任意可用性要求,尽管临时虚拟机不可用,然后当有多余资源可用时,利用这些虚拟机机会性地执行批处理作业。当云平台撤销临时虚拟机时,TR - Kubernetes依靠现有的Kubernetes功能从内部撤销批处理作业的资源,以维持交互式服务的可用性要求。我们表明,TR - Kubernetes对Kubernetes的扩展要求极少,并且在使用临时虚拟机(与按需虚拟机相比)时,能够降低在亚马逊EC2上具有代表性的交互式/批处理工作负载的成本(降低53%)并提高可用性(达到99.999%)。
Container Orchestration Platforms (COPs), such as Kubernetes, are increasingly used to manage large-scale clusters by automating resource allocation between applications encapsulated in containers. Increasingly, the resources underlying COPs are virtual machines (VMs) dynamically acquired from cloud platforms. COPs may choose from many different types of VMs offered by cloud platforms, which differ in their cost, performance, and availability. While transient VMs cost significantly less than on-demand VMs, platforms may revoke them at any time, causing them to become unavailable. While transient VMs' price is attractive, their unreliability is a problem for COPs designed to support mixed workloads composed of, not only delay-tolerant batch jobs, but also long-lived interactive services with high availability requirements. To address the problem, we design TR-Kubernetes, a COP that optimizes the cost of executing mixed interactive and batch workloads on cloud platforms using transient VMs. To do so, TR-Kubernetes enforces arbitrary availability requirements specified by interactive services despite transient VM unavailability by acquiring many more transient VMs than necessary most of the time, which it then leverages to opportunistically execute batch jobs when excess resources are available. When cloud platforms revoke transient VMs, TR-Kubernetes relies on existing Kubernetes functions to internally revoke resources from batch jobs to maintain interactive services' availability requirements. We show that TR-Kubernetes requires minimal extensions to Kubernetes, and is capable of lowering the cost (by 53%) and improving the availability (99.999%) of a representative interactive/batch workload on Amazon EC2 when using transient compared to on-demand VMs.