SplitServe: Efficiently Splitting Apache Spark Jobs Across FaaS and IaaS

SplitServe: Efficiently Splitting Apache Spark Jobs Across FaaS and IaaS
复制标题

DOI:
10.1145/3423211.3425695
复制
发表时间:
2020-12
期刊:
Proceedings of the 21st International Middleware Conference
影响因子:
--
通讯作者:
Aman Jain;A. F. Baarzi;G. Kesidis;B. Urgaonkar;Nader Alfares;M. Kandemir
Aman Jain;A. F. Baarzi;G. Kesidis;B. Urgaonkar;Nader Alfares;M. Kandemir
中科院分区:
其他
文献类型:
--
作者:
Aman Jain;A. F. Baarzi;G. Kesidis;B. Urgaonkar;Nader Alfares;M. Kandemir

文献摘要

相似文献

Amazon Lambdas和其他云功能(CF)的启动延迟更低,定价比虚拟机(VM)更精细,因此被认为是处理简单、无状态工作负载中意外峰值的理想选择。然而,目前还不清楚CFS在自动扩展复杂工作负载方面是否同样有效,这些工作负载涉及跨分布式应用程序组件的大量状态传输。我们发现,通过仔细的设计,即使对于复杂的工作负载,当前可用的CFS也确实有用。为了演示这一点,我们设计并实现了SplitServe,这是对ApacheSpark的增强。如果现有VM上没有足够的执行器可用于新到达的延迟敏感型作业,SplitServe能够使用CFS快速弥补VM中的这一不足,从而避免新请求的VM的启动延迟。如果在性能或成本方面需要,当新请求的VM或现有VM上的执行器可用时,SplitServe能够将正在进行的工作从CFS转移到它们。我们使用四种不同的工作负载对SplitServe进行的实验评估(无论是在基于VM的执行器和CFS的混合体上,还是仅在CFS上)表明,与仅基于VM的自动伸缩相比,SplitServe的执行时间最多可缩短(A)55%(对于具有少量到中等洗牌的工作负载)和(B)31%(对于具有大量洗牌的工作负载)。
Due to their lower startup latencies and finer-grain pricing than virtual machines (VMs), Amazon Lambdas and other cloud functions (CFs) have been identified as ideal candidates for handling unexpected spikes in simple, stateless workloads. However, it is not immediately clear if CFs would be similarly effective in autoscaling complex workloads involving significant state transfer across distributed application components. We have found that, through careful design, currently available CFs can indeed be useful even for complex workloads. To demonstrate this, we design and implement SplitServe, an enhancement of Apache Spark. If not enough executors on existing VMs are available for a newly arriving latency-sensitive job, SplitServe is able to use CFs to quickly bridge this shortfall in VMs, so avoiding the startup latencies of newly requested VMs. If desirable in terms of performance or cost, when newly requested VMs, or executors on existing VMs, do become available, SplitServe is able to move ongoing work from CFs to them. Our experimental evaluation of SplitServe using four different workloads (either on a mixture of VM-based executors and CFs or just CFs) shows that it improves execution time by up to (a) 55% for workloads with small to modest amount of shuffling, and (b) 31% in workloads with large amounts of shuffling, when compared to only VM-based autoscaling.