Multiverse: Dynamic VM Provisioning for Virtualized High Performance Computing Clusters

Multiverse: Dynamic VM Provisioning for Virtualized High Performance Computing Clusters
复制标题

DOI:
10.1109/ccgrid49817.2020.00-80
复制
发表时间:
2020-05
期刊:
2020 20th IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGRID)
影响因子:
--
通讯作者:
Jashwant Raj Gunasekaran;Michael Cui;P. Thinakaran;Josh Simons;M. Kandemir;C. Das
Jashwant Raj Gunasekaran;Michael Cui;P. Thinakaran;Josh Simons;M. Kandemir;C. Das
中科院分区:
其他
文献类型:
--
作者:
Jashwant Raj Gunasekaran;Michael Cui;P. Thinakaran;Josh Simons;M. Kandemir;C. Das

文献摘要

相似文献

传统上,HPC工作负载部署在裸机群集中;但虚拟化的进步引领了这些工作负载部署到虚拟群集中的道路。然而,由于传统的HPC调度器和VM管理程序(资源管理层)之间缺乏协调,HPC集群管理员/提供商在大规模的资源弹性和虚拟机(VM)配置方面仍然面临挑战。这种缺乏交互会导致集群利用率和作业完成吞吐量较低。此外,VM资源调配延迟直接影响集群中作业的整体性能。因此,需要有效地配置虚拟HPC集群,从而以最小的配置开销最大限度地利用物理硬件。为此,我们提出了一种VM配置框架MultiVerse,它可以通过集成HPC调度器和VM资源管理器来动态地为虚拟HPC集群中的传入作业生成VM。我们已经在SLurm调度程序和vSphere VM资源管理器上实施了此框架。为了减少VM资源调配开销,我们使用即时克隆,与必须从头启动新VM的完全VM克隆相比,它与父VM共享磁盘和内存。使用真实HPC工作负载进行的测量表明,就VM资源调配时间而言,即时克隆比完全克隆快2.5倍。此外,与突发作业到达情况下的完全克隆相比,它可将资源利用率提高高达40%,将群集吞吐量提高高达1.5倍。
Traditionally, HPC workloads have been deployed in bare-metal clusters; but the advances in virtualization have led the pathway for these workloads to be deployed in virtualized clusters. However, HPC cluster administrators/providers still face challenges in terms of resource elasticity and virtual machine (VM) provisioning at large-scale, due to the lack of coordination between a traditional HPC scheduler and the VM hypervisor (resource management layer). This lack of interaction leads to low cluster utilization and job completion throughput. Furthermore, the VM provisioning delays directly impact the overall performance of jobs in the cluster. Hence, there is a need for effectively provisioning virtualized HPC clusters, which can best-utilize the physical hardware with minimal provisioning overheads.Towards this, we propose Multiverse, a VM provisioning framework, which can dynamically spawn VMs for incoming jobs in a virtualized HPC cluster, by integrating the HPC scheduler along with VM resource manager. We have implemented this framework on the Slurm scheduler along with the vSphere VM resource manager. In order to reduce the VM provisioning overheads, we use instant cloning which shares both the disk and memory with the parent VM, when compared to full VM cloning which has to boot-up a new VM from scratch. Measurements with real-world HPC workloads demonstrate that, instant cloning is 2.5× faster than full cloning in terms of VM provisioning time. Further, it improves resource utilization by up to 40%, and cluster throughput by up to 1.5×, when compared to full clone for bursty job arrival scenarios.