Efficient warp execution in presence of divergence with collaborative context collection

Efficient warp execution in presence of divergence with collaborative context collection
复制标题

通过协作上下文收集在存在分歧的情况下高效执行扭曲

DOI:
10.1145/2830772.2830796
复制
发表时间:
2015
期刊:
2015 48th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
L. Bhuyan
L. Bhuyan
中科院分区:
--
文献类型:
--
作者:
Farzad Khorasani;Rajiv Gupta;L. Bhuyan

文献摘要

被引文献

相似文献

GPU的SIMD体系结构是一项双边的剑,一方面是控制流动差异,它提供了一个高性能但有能力的平台,以通过大规模的并行性加速应用。在所有不同的执行路径中,Warp的锁定遍历。收集(CCC)在面对线程差异时提高了经常执行效率共享记忆,并且只有在可行的扭曲车道的完美利用率时才能恢复它们。线程差异。使用合成程序的方案。 1.69×(最大3.08×)。
GPU's SIMD architecture is a double-edged sword confronting parallel tasks with control flow divergence. On the one hand, it provides a high performance yet power-efficient platform to accelerate applications via massive parallelism; however, on the other hand, irregularities induce inefficiencies due to the warp's lockstep traversal of all diverging execution paths. In this work, we present a software (compiler) technique named Collaborative Context Collection (CCC) that increases the warp execution efficiency when faced with thread divergence incurred either by different intra-warp task assignment or by intra-warp load imbalance. CCC collects the relevant registers of divergent threads in a warp-specific stack allocated in the fast shared memory, and restores them only when the perfect utilization of warp lanes becomes feasible. We propose code transformations to enable applicability of CCC to variety of program segments with thread divergence. We also introduce optimizations to reduce the cost of CCC and to avoid device occupancy limitation or memory divergence. We have developed a framework that automates application of CCC to CUDA generated intermediate PTX code. We evaluated CCC on real-world applications and multiple scenarios using synthetic programs. CCC improves the warp execution efficiency of real-world benchmarks by up to 56% and achieves an average speedup of 1.69× (maximum 3.08×).