CASH: A Credit Aware Scheduling for Public Cloud Platforms

CASH: A Credit Aware Scheduling for Public Cloud Platforms
复制标题

DOI:
10.1109/ccgrid51090.2021.00032
复制
发表时间:
2020-09
期刊:
2021 IEEE/ACM 21st International Symposium on Cluster, Cloud and Internet Computing (CCGrid)
影响因子:
--
通讯作者:
Aakash Sharma;Saravanan Dhakshinamurthy;G. Kesidis;C. Das
Aakash Sharma;Saravanan Dhakshinamurthy;G. Kesidis;C. Das
中科院分区:
其他
文献类型:
--
作者:
Aakash Sharma;Saravanan Dhakshinamurthy;G. Kesidis;C. Das

文献摘要

相似文献

公共云租户专门使用了分布式数据处理框架,例如Hadoop,Tez,Spark和Flink,用于执行大规模数据分析应用程序在各个领域的大规模数据分析应用程序,包括但不限于内容管理,财务部门,医疗保健等。这些框架将作业切成了许多较小的任务,然后由工作调整器执行,然后由一项工作调整器执行,这些任务是由一项工作时间表上的多人计算机计算。在做出计划决策时,这些框架中使用的最先进的调度程序假设硬件资源(例如CPU,磁盘I/O和网络I/O)提供固定的服务率。但是,在公共云环境中,其中许多资源与可爆破的服务率有关。更具体地说,资源提供了保证的基线服务率,可以通过花费累积的爆发信用来选择超过基线率的选择。对于这种潜在的硬件破裂性,调度程序倾向于做出次优的任务放置决策,从而对工作完成时间产生不利影响,从而导致更高的部署成本。在本文中,我们提出了兑现,这是一份爆发的信贷意识调整器,这是对公共云云中个人硬件资源相关的爆发信用。通过粗糙的任务注释,描绘了单个任务的信用要求突破,并动态监视基础资源的信用,现金执行最佳任务安置决策。我们在纱线,hadoop和tez上原型现金,并使用批处理和流式工作负载进行了广泛的评估。我们使用现金的实验结果表明,与Amazon EMR这样的自我管理产品相比,与AWS T3这样的基于CPU-CREDIT的实例(例如AWS T3)是可行的成本效益替代方案。此外,我们证明现金可以在大型Hive数据库上加速流式传输SQL查询,高达39.4%,从而使公共云成本节省高达22%。
Distributed data processing frameworks such as Hadoop, Tez, Spark, and Flink are exclusively used by public cloud tenants for executing large scale data analytics applications in various domains including but not limited to content management, financial sector, healthcare etc. These frameworks slice a job into a number of smaller tasks, which are then executed by a job scheduler on a multi-node compute cluster. While making scheduling decisions, the State-of-art schedulers employed in these frameworks assume hardware resources such as CPU, disk I/O and network I/O to offer a fixed service rate. However, in a public cloud environment, many of these resources are associated with burstable service rates. More specifically, the resources offer a guaranteed baseline service rate with an option to burst above their baseline rate by expending accumulated burst credits. Being unaware about this underlying hardware burstability, schedulers tend to make sub-optimal task placement decisions, thereby adversely affecting the job completion times, leading to higher deployment costs.In this paper, we propose CASH, a burst credit aware scheduler, which is cognizant about the burst credits associated with the individual hardware resources in the public cloud cluster. Through coarse grained task annotations depicting the burst credit demand of individual tasks and dynamically monitoring the credits for the underlying resources, CASH performs optimal task placement decisions. We prototype CASH on YARN, Hadoop, and Tez, and extensively evaluate it using both batch and streaming workloads. Our experimental results with CASH show CPU-credit based instances, like AWS T3, are a viable cost effective alternative when compared to self-managed offerings like Amazon EMR, for running large scale batch workloads. Furthermore, we demonstrate that CASH can accelerate streaming SQL queries on a large Hive database by up to 39.4% , leading to public cloud cost savings by up to 22%.