A Hypervisor for Shared-Memory FPGA Platforms

A Hypervisor for Shared-Memory FPGA Platforms
复制标题

适用于共享内存 FPGA 平台的虚拟机管理程序

DOI:
10.1145/3373376.3378482
复制
发表时间:
2020
期刊:
Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Baris Kasikci
Baris Kasikci
中科院分区:
--
文献类型:
--
作者:
Jiacheng Ma;Gefei Zuo;Kevin Loughlin;Xiao;Yanqiang Liu;Abel Mulugeta Eneyew;Zhengwei Qi;Baris Kasikci

文献摘要

参考文献

被引文献

相似文献

云提供商广泛部署FPGA作为用于客户使用的特定应用程序的加速器。这些提供商试图通过虚拟化将其FPGA在客户之间多样化,从而降低运行成本。不幸的是,大多数虚拟化支持仅限于FPGA,该FPGA暴露了一个限制性的,以宿主为中心的编程模型,在该模型中,加速器无法发布直接内存访问(DMA)。以主机为中心的模型为展示指针追逐的工作负载带来了高运行时开销。因此,FPGA开始支持共享的内存编程模型,在该模型中,加速器可以发行DMA。但是,对共享内存FPGA的虚拟化支持是有限的。本文介绍了Optimus,这是支持可扩展共享内存FPGA虚拟化的第一个管理程序。 Optimus提供空间多路复用和时间多路复用,以在FPGA上提供每个加速器的有效和灵活共享。要以高时钟频率共享FPGA-CPU互连,Optimus实现了多路复用器树。为了隔离每个客人的地址空间,Optimus将页面表切片的技术作为硬件软件共同设计。为了支持先发制的时间多路复用,Optimus提供了加速器的抢先接口。我们表明,Optimus支持单个FPGA上的八个物理加速器,并将十二个现实世界基准的总吞吐量提高了1.98x-7x。
Cloud providers widely deploy FPGAs as application-specific accelerators for customer use. These providers seek to multiplex their FPGAs among customers via virtualization, thereby reducing running costs. Unfortunately, most virtualization support is confined to FPGAs that expose a restrictive, host-centric programming model in which accelerators cannot issue direct memory accesses (DMAs). The host-centric model incurs high runtime overhead for workloads that exhibit pointer chasing. Thus, FPGAs are beginning to support a shared-memory programming model in which accelerators can issue DMAs. However, virtualization support for shared-memory FPGAs is limited. This paper presents Optimus, the first hypervisor that supports scalable shared-memory FPGA virtualization. Optimus offers both spatial multiplexing and temporal multiplexing to provide efficient and flexible sharing of each accelerator on an FPGA. To share the FPGA-CPU interconnect at a high clock frequency, Optimus implements a multiplexer tree. To isolate each guest's address space, Optimus introduces the technique of page table slicing as a hardware-software co-design. To support preemptive temporal multiplexing, Optimus provides an accelerator preemption interface. We show that Optimus supports eight physical accelerators on a single FPGA and improves the aggregate throughput of twelve real-world benchmarks by 1.98x-7x.
DOI: 10.1145/3297858.3304010
发表时间: 2019
期刊: Proceedings of the Twenty-Fourth International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子: --
作者:
Schkufza, Eric;Wei, Michael;Rossbach, Christopher J.
通讯作者: Rossbach, Christopher J.