MASK: Redesigning the GPU Memory Hierarchy to Support Multi-Application Concurrency

MASK: Redesigning the GPU Memory Hierarchy to Support Multi-Application Concurrency
复制标题

DOI:
10.1145/3173162.3173169
复制
发表时间:
2018-03
期刊:
Proceedings of the Twenty-Third International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Rachata Ausavarungnirun;Vance Miller;Joshua Landgraf;Saugata Ghose;Jayneel Gandhi;Adwait Jog;C. Rossbach;O. Mutlu
Rachata Ausavarungnirun;Vance Miller;Joshua Landgraf;Saugata Ghose;Jayneel Gandhi;Adwait Jog;C. Rossbach;O. Mutlu
中科院分区:
其他
文献类型:
--
作者:
Rachata Ausavarungnirun;Vance Miller;Joshua Landgraf;Saugata Ghose;Jayneel Gandhi;Adwait Jog;C. Rossbach;O. Mutlu

文献摘要

被引文献

相似文献

图形处理单元(GPU)利用大量的线程级并行性来提供高指令吞吐量并有效地隐藏长延迟停顿。由此产生的高吞吐量,沿着持续的可编程性改进,使GPU成为许多领域中必不可少的计算资源。来自不同领域的应用程序对GPU的计算和内存需求可能有很大不同。在大规模计算环境中,为了有效地适应这种广泛的需求而不使GPU资源未得到充分利用,多个应用程序可以共享单个GPU,类似于多个应用程序如何在CPU上并发执行。多应用程序并发需要硬件和软件中的多个支持机制。其中一个关键机制是虚拟内存,它管理和保护每个应用程序的地址空间。然而,现代GPU缺乏对CPU中可用的多应用程序并发的广泛支持,因此当被多个应用程序共享时,会遭受高性能开销,正如我们所展示的那样。我们进行了详细的分析,多应用程序并发支持的限制损害GPU性能最多。我们发现,性能不佳主要是由于现代GPU中采用的虚拟内存机制。特别是,较差的地址转换性能是高效GPU共享的关键障碍。最先进的地址转换机制是为单个应用程序执行而设计的,当多个应用程序在空间上共享GPU时,会遇到严重的应用程序间干扰。这种争用导致共享转换后备缓冲区(TLB)中频繁的未命中,其中单个未命中可能导致数百个线程的长延迟暂停。因此,GPU通常无法调度足够的线程来成功隐藏停顿,这会降低系统吞吐量并成为一个首要的性能问题。根据我们的分析,我们提出了MASK,这是一种新的GPU框架,为多个应用程序的并发执行提供低开销虚拟内存支持。MASK由三种新颖的地址转换感知的缓存和内存管理机制组成,它们共同工作以大大减少地址转换的开销:(1)基于令牌的技术以减少TLB争用,(2)旁路机制以提高缓存地址转换的有效性,以及(3)应用感知的内存调度方案以减少地址转换和数据请求之间的干扰。我们的评估表明,MASK恢复了TLB争用造成的大部分吞吐量损失。相对于最先进的GPU TLB,MASK将系统吞吐量提高了57.8%,将IPC吞吐量提高了43.4%,并将应用级不公平性降低了22.4%。MASK的系统吞吐量在理想GPU系统的23.2%以内,没有地址转换开销。
Graphics Processing Units (GPUs) exploit large amounts of threadlevel parallelism to provide high instruction throughput and to efficiently hide long-latency stalls. The resulting high throughput, along with continued programmability improvements, have made GPUs an essential computational resource in many domains. Applications from different domains can have vastly different compute and memory demands on the GPU. In a large-scale computing environment, to efficiently accommodate such wide-ranging demands without leaving GPU resources underutilized, multiple applications can share a single GPU, akin to how multiple applications execute concurrently on a CPU. Multi-application concurrency requires several support mechanisms in both hardware and software. One such key mechanism is virtual memory, which manages and protects the address space of each application. However, modern GPUs lack the extensive support for multi-application concurrency available in CPUs, and as a result suffer from high performance overheads when shared by multiple applications, as we demonstrate. We perform a detailed analysis of which multi-application concurrency support limitations hurt GPU performance the most. We find that the poor performance is largely a result of the virtual memory mechanisms employed in modern GPUs. In particular, poor address translation performance is a key obstacle to efficient GPU sharing. State-of-the-art address translation mechanisms, which were designed for single-application execution, experience significant inter-application interference when multiple applications spatially share the GPU. This contention leads to frequent misses in the shared translation lookaside buffer (TLB), where a single miss can induce long-latency stalls for hundreds of threads. As a result, the GPU often cannot schedule enough threads to successfully hide the stalls, which diminishes system throughput and becomes a first-order performance concern. Based on our analysis, we propose MASK, a new GPU framework that provides low-overhead virtual memory support for the concurrent execution of multiple applications. MASK consists of three novel address-translation-aware cache and memory management mechanisms that work together to largely reduce the overhead of address translation: (1) a token-based technique to reduce TLB contention, (2) a bypassing mechanism to improve the effectiveness of cached address translations, and (3) an application-aware memory scheduling scheme to reduce the interference between address translation and data requests. Our evaluations show that MASK restores much of the throughput lost to TLB contention. Relative to a state-of-the-art GPU TLB, MASK improves system throughput by 57.8%, improves IPC throughput by 43.4%, and reduces applicationlevel unfairness by 22.4%. MASK's system throughput is within 23.2% of an ideal GPU system with no address translation overhead.