Ninja Migration: An Interconnect-Transparent Migration for Heterogeneous Data Centers

Ninja Migration: An Interconnect-Transparent Migration for Heterogeneous Data Centers
复制标题

DOI:
10.1109/ipdpsw.2013.114
复制
发表时间:
2013-05
期刊:
2013 IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum
影响因子:
--
通讯作者:
Ryousei Takano;H. Nakada;Takahiro Hirofuchi;Yoshio Tanaka;T. Kudoh
Ryousei Takano;H. Nakada;Takahiro Hirofuchi;Yoshio Tanaka;T. Kudoh
中科院分区:
其他
文献类型:
--
作者:
Ryousei Takano;H. Nakada;Takahiro Hirofuchi;Yoshio Tanaka;T. Kudoh

文献摘要

相似文献

虚拟机 (VM) 迁移对于提高云计算环境中的灵活性和可维护性非常有用。然而,底层软件和硬件(包括CPU和互连架构)的异构性使得虚拟机迁移变得困难。此外,虚拟机监控(VMM)-旁路I/O技术显着降低了I/O虚拟化的开销,也使虚拟机迁移变得不可能。因此,分配给 Infiniband 设备的 VM 无法迁移到以太网计算机,反之亦然。如果我们克服上述障碍,我们就可以增加虚拟机迁移的潜力和可能性。在本文中,我们提出了一种互连透明的迁移机制,可以在配备不同互连设备的数据中心之间同时迁移多个共置虚拟机。我们所提出的机制(称为 Ninja 迁移)的实现是通过 VMM 和来宾操作系统上的 MPI 运行时系统之间的合作来实现的。我们使用所提出的机制演示了高性能计算工作负载的回退和恢复操作。我们已经确认:1)所提出的机制在正常操作期间没有性能开销,2)在分布式虚拟机上运行的 MPI 进程可以在 Infiniband 集群和以太网集群之间迁移,而无需重新启动进程。
A virtual machine (VM) migration is useful for improving flexibility and maintainability in cloud computing environments. However, the heterogeneity of the underlying software and hardware, including CPU and interconnect architectures, makes it hard to migrate a VM. In addition, VM monitor~(VMM)-bypass I/O technologies, which significantly reduce the overhead of I/O virtualization, also make VM migration impossible. Therefore, a VM assigned to an Infiniband device cannot migrate to an Ethernet machine, and vice versa. If we overcome the above barriers, we can increase the potential and possibilities of VM migration. In this paper, we propose an interconnect-transparent migration mechanism to simultaneously migrate multiple co-located VMs between data centers equipped with different interconnect devices. Our implementation of the proposed mechanism, called Ninja migration, is achieved by cooperation between a VMM and an MPI runtime system on the guest OS. We demonstrate fallback and recovery operations on a high performance computing workload using the proposed mechanism. We have confirmed that 1) the proposed mechanism has no performance overhead during normal operations, and 2) MPI processes running on distributed VMs can migrate between an Infiniband cluster and an Ethernet cluster without restarting the processes.