GPU Performance vs. Thread-Level Parallelism

GPU Performance vs. Thread-Level Parallelism
复制标题

DOI:
10.1145/3177964
复制
发表时间:
2018-03
期刊:
ACM Transactions on Architecture and Code Optimization (TACO)
影响因子:
--
通讯作者:
Zhen Lin;Mike Mantor;Huiyang Zhou
Zhen Lin;Mike Mantor;Huiyang Zhou
中科院分区:
其他
文献类型:
--
作者:
Zhen Lin;Mike Mantor;Huiyang Zhou

文献摘要

被引文献

相似文献

图形处理单元(GPU)利用大规模线程级并行(TLP)来实现高计算吞吐量并隐藏长内存延迟。然而,最近的研究表明,GPU的性能并不与GPU占用率或GPU支持的TLP程度成比例,特别是对于内存密集型工作负载。目前的理解指向L1 D缓存争用或片外内存带宽。在本文中,我们从各种GPU组件(包括片外DRAM、多级缓存以及L1 D缓存和L2分区之间的互连)的吞吐量利用率的角度进行了一种新颖的可扩展性分析。我们表明,互连带宽是GPU性能可扩展性的关键界限。对于在特定资源上没有饱和吞吐量利用率的应用程序,它们的性能随着TLP的增加而很好地扩展。为了有效地提高TLP,这样的应用程序,我们提出了一个快速的上下文切换方法。当线程束/线程块(TB)被长延迟操作停止时,线程束/TB的上下文被溢出以备用片上资源,使得可以启动新的线程束/TB。当另一个warp/TB完成或切换出时,切换出的warp/TB会切换回来。通过这种细粒度的快速上下文切换,可以支持更高的TLP,而无需增加寄存器文件等关键资源的大小。我们的实验表明,对于一组吞吐量利用率不饱和的应用程序,性能可以提高高达47%,几何平均值为22%。与最先进的TLP改进方案相比,我们提出的方案平均性能提高了12%,对于不饱和基准测试提高了16%。
Graphics Processing Units (GPUs) leverage massive thread-level parallelism (TLP) to achieve high computation throughput and hide long memory latency. However, recent studies have shown that the GPU performance does not scale with the GPU occupancy or the degrees of TLP that a GPU supports, especially for memory-intensive workloads. The current understanding points to L1 D-cache contention or off-chip memory bandwidth. In this article, we perform a novel scalability analysis from the perspective of throughput utilization of various GPU components, including off-chip DRAM, multiple levels of caches, and the interconnect between L1 D-caches and L2 partitions. We show that the interconnect bandwidth is a critical bound for GPU performance scalability. For the applications that do not have saturated throughput utilization on a particular resource, their performance scales well with increased TLP. To improve TLP for such applications efficiently, we propose a fast context switching approach. When a warp/thread block (TB) is stalled by a long latency operation, the context of the warp/TB is spilled to spare on-chip resource so that a new warp/TB can be launched. The switched-out warp/TB is switched back when another warp/TB is completed or switched out. With this fine-grain fast context switching, higher TLP can be supported without increasing the sizes of critical resources like the register file. Our experiment shows that the performance can be improved by up to 47% and a geometric mean of 22% for a set of applications with unsaturated throughput utilization. Compared to the state-of-the-art TLP improvement scheme, our proposed scheme achieves 12% higher performance on average and 16% for unsaturated benchmarks.