GPUvm: Why Not Virtualizing GPUs at the Hypervisor?

GPUvm: Why Not Virtualizing GPUs at the Hypervisor?
复制标题

DOI:
--
复制
发表时间:
2014-06
期刊:
影响因子:
3.1
通讯作者:
Yusuke Suzuki;S. Kato;H. Yamada;K. Kono
Yusuke Suzuki;S. Kato;H. Yamada;K. Kono
中科院分区:
化学3区
文献类型:
--
作者:
Yusuke Suzuki;S. Kato;H. Yamada;K. Kono

文献摘要

被引文献

相似文献

图形处理单元(GPU)为计算密集型数据并行应用提供了数量级的加速。然而,企业和云计算领域需要多个客户端的资源隔离,对GPU技术的访问很差。这是由于缺乏操作系统(OS)支持以可靠的方式虚拟化GPU。为了使GPU成为更成熟的系统公民,我们提出了一个开放的GPU虚拟化架构,特别强调了Xen虚拟机管理程序。我们提供全虚拟化和半虚拟化的设计和实施,包括降低GPU虚拟化开销的优化技术。我们使用相关商品GPU进行的详细实验表明,GPU半虚拟化的优化性能比直通和本机方法慢两到三倍,而全虚拟化由于增加了内存映射I/O操作而表现出不同规模的开销。我们还证明了多个虚拟机之间的GPU资源的粗粒度公平性可以通过GPU调度来实现;由于非抢占式GPU工作负载的性质,细粒度公平性需要进一步的架构支持。
Graphics processing units (GPUs) provide orders-of-magnitude speedup for compute-intensive data-parallel applications. However, enterprise and cloud computing domains, where resource isolation of multiple clients is required, have poor access to GPU technology. This is due to lack of operating system (OS) support for virtualizing GPUs in a reliable manner. To make GPUs more mature system citizens, we present an open architecture of GPU virtualization with a particular emphasis on the Xen hypervisor. We provide design and implementation of full- and para-virtualization, including optimization techniques to reduce overhead of GPU virtualization. Our detailed experiments using a relevant commodity GPU show that the optimized performance of GPU para-virtualization is yet two or three times slower than that of pass-through and native approaches, whereas full-virtualization exhibits a different scale of overhead due to increased memory-mapped I/O operations. We also demonstrate that coarse-grained fairness on GPU resources among multiple virtual machines can be achieved by GPU scheduling; finer-grained fairness needs further architectural support by the nature of non-preemptive GPU workload.