CAWS: Criticality-aware warp scheduling for GPGPU workloads

CAWS: Criticality-aware warp scheduling for GPGPU workloads
复制标题

DOI:
10.1145/2628071.2628107
复制
发表时间:
2014-08
期刊:
2014 23rd International Conference on Parallel Architecture and Compilation (PACT)
影响因子:
--
通讯作者:
Shin-Ying Lee;Carole-Jean Wu
Shin-Ying Lee;Carole-Jean Wu
中科院分区:
其他
文献类型:
--
作者:
Shin-Ying Lee;Carole-Jean Wu

文献摘要

被引文献

相似文献

执行快速上下文切换和大规模多线程的能力是现代GPU架构的强项,它已成为传统芯片多处理器的有效替代品,用于并行工作负载。这种架构的主要优点之一是其延迟隐藏能力。然而,GPU的延迟隐藏的功效在GPGPU应用程序中差异很大。为了研究这一点,本文首先提出了一种新的算法,分析GPGPU应用程序的执行行为。我们描述了由各种管道危险、内存访问、同步原语和warp调度器引起的延迟。我们的研究结果表明,目前的循环线程调度器工作良好,在重叠的各种延迟失速与其他可用的线程的执行只有少数GPGPU应用程序。对于其他应用程序,调度程序无法有效地隐藏过多的延迟延迟。随着延迟特性的洞察力,我们观察到一个显着的执行时间差异的线程在同一个线程块内,这会导致次优性能,称为线程的关键性问题。为了解决翘曲关键性问题,我们设计了一个家庭的关键性感知翘曲调度(CAWS)的政策,调度的关键翘曲(S)比其他翘曲更频繁。在广度优先搜索、B+树搜索、两点角相关函数和K-means聚类等方面的实验结果表明,在Oracle知识的指导下,本文提出的最佳调度策略可以使GPGPU应用程序的性能平均提高17%.使用我们设计的关键度预测器,各种调度策略可以在广度优先搜索上提高10-21%的性能。据我们所知,这是第一篇描述warp关键性并探索GPGPU工作负载的不同关键性感知warp调度策略的论文。
The ability to perform fast context-switching and massive multi-threading is the forte of modern GPU architectures, which have emerged as an efficient alternative to traditional chip-multiprocessors for parallel workloads. One of the main benefits of such architecture is its latency-hiding capability. However, the efficacy of GPU's latency-hiding varies significantly across GPGPU applications. To investigate this, this paper first proposes a new algorithm that profiles execution behavior of GPGPU applications. We characterize latencies caused by various pipeline hazards, memory accesses, synchronization primitives, and the warp scheduler. Our results show that the current round-robin warp scheduler works well in overlapping various latency stalls with the execution of other available warps for only a few GPGPU applications. For other applications, there is an excessive latency stall that cannot be hidden by the scheduler effectively. With the latency characterization insight, we observe a significant execution time disparity for warps within the same thread block, which causes suboptimal performance, called the warp criticality problem. To tackle the warp criticality problem, we design a family of criticality-aware warp scheduling (CAWS) policies by scheduling the critical warp(s) more frequently than other warps. Our results on the breadth-first-search, B+tree search, two point angular correlation function, and K-means clustering show that, with oracle knowledge of warp criticality, our best-performing scheduling policy can improve GPGPU applications' performance by 17% on average. With our designed criticality predictor, the various scheduling policies can improve performance by 10–21% on breadth-first-search. To our knowledge, this is the first paper to characterize warp criticality and explore different criticality-aware warp scheduling policies for GPGPU workloads.