Benchmarking the Cost of Thread Divergence in CUDA

Benchmarking the Cost of Thread Divergence in CUDA
复制标题

CUDA 中线程分歧成本的基准测试

DOI:
--
复制
发表时间:
2015
期刊:
Parallel Processing and Applied Mathematics
影响因子:
--
通讯作者:
A. Strzelecki
A. Strzelecki
中科院分区:
--
文献类型:
--
作者:
P. Bialas;A. Strzelecki

文献摘要

被引文献

相似文献

所有现代处理器都包含一组向量说明。尽管这给了性能巨大的提升,但它需要一个可以利用此类说明的矢量代码。由于在实践中很难实现理想的矢量化,因此必须决定何时可以将不同的指示应用于向量操作数的不同元素。这在隐式矢量化中尤为重要,就像NVIDIA CUDA单个指令多个线程(SIMT)模型中一样,其中矢量化细节隐藏在程序员中。为了评估未完全矢量化的代码所产生的成本,我们开发了一个微基准测试,该基准测量了CUDA线程散射模型的特征,这些模型是针对循环性能的不同体系结构。
All modern processors include a set of vector instructions. While this gives a tremendous boost to the performance, it requires a vectorized code that can take advantage of such instructions. As an ideal vectorization is hard to achieve in practice, one has to decide when different instructions may be applied to different elements of the vector operand. This is especially important in implicit vectorization as in NVIDIA CUDA Single Instruction Multiple Threads (SIMT) model, where the vectorization details are hidden from the programmer. In order to assess the costs incurred by incompletely vectorized code, we have developed a micro-benchmark that measures the characteristics of the CUDA thread divergence model on different architectures focusing on the loops performance.