uBench: exposing the impact of CUDA block geometry in terms of performance

uBench: exposing the impact of CUDA block geometry in terms of performance
复制标题

DOI:
10.1007/s11227-013-0921-z
复制
发表时间:
2013-04
期刊:
The Journal of Supercomputing
影响因子:
--
通讯作者:
Yuri Torres;Arturo González-Escribano;D. Ferraris
Yuri Torres;Arturo González-Escribano;D. Ferraris
中科院分区:
其他
文献类型:
--
作者:
Yuri Torres;Arturo González-Escribano;D. Ferraris

文献摘要

被引文献

相似文献

当为任何CUDA架构编写并行问题时,线程块大小和形状的选择是最重要的用户决策之一。原因是线程块几何结构对程序的全局性能有很大的影响。不幸的是,程序员没有足够的信息之间的微妙的相互作用,这种选择的参数和底层hardware.This本文介绍了uBench,一套完整的微benchmark,以探讨性能的影响(1)线程块的几何选择标准,(2)GPU的硬件资源和配置。每个微基准测试都被设计得尽可能简单,专注于从硬件和线程块参数选择中获得的单一效果。作为该基准测试套件功能的一个例子,本文展示了费米和开普勒架构的实验评估和比较。我们的研究表明,尽管开普勒介绍了新的硬件细节,块几何选择标准的基本原则是相似的两种架构。
The choice of thread-block size and shape is one of the most important user decisions when a parallel problem is written for any CUDA architecture. The reason is that thread-block geometry has a significant impact on the global performance of the program. Unfortunately, the programmer has not enough information about the subtle interactions between this choice of parameters and the underlying hardware.This paper presents uBench, a complete suite of micro-benchmarks, in order to explore the impact on performance of (1) the thread-block geometry choice criteria, and (2) the GPU hardware resources and configurations. Each micro-benchmark has been designed to be as simple as possible to focus on a single effect derived from the hardware and thread-block parameter choice.As an example of the capabilities of this benchmark suite, this paper shows an experimental evaluation and comparison of Fermi and Kepler architectures. Our study reveals that, in spite of the new hardware details introduced by Kepler, the principles underlying the block geometry selection criteria are similar for both architectures.