Understanding the impact of CUDA tuning techniques for Fermi

Understanding the impact of CUDA tuning techniques for Fermi
复制标题

DOI:
10.1109/hpcsim.2011.5999886
复制
发表时间:
2011-07
期刊:
2011 International Conference on High Performance Computing & Simulation
影响因子:
--
通讯作者:
Yuri Torres;Arturo González-Escribano;D. Ferraris
Yuri Torres;Arturo González-Escribano;D. Ferraris
中科院分区:
其他
文献类型:
--
作者:
Yuri Torres;Arturo González-Escribano;D. Ferraris

文献摘要

被引文献

相似文献

虽然 NVIDIA CUDA 程序的正确性很容易实现,但利用 GPU 功能来获得尽可能最佳的性能对于 CUDA 经验丰富的程序员来说是一项任务。典型的代码调整策略,例如为线程块选择适当的大小和形状、编程良好的合并或最大化占用率,是相互依赖的。此外,选择还取决于底层架构细节以及设计解决方案的全局内存访问模式。例如,通常选择线程块的大小和形状以方便编码(例如正方形),同时最大化多处理器的占用。然而,这种简单的选择通常不能提供最佳的性能结果。在本文中,我们讨论了线程块的大小和形状、占用、全局内存访问模式以及其他 Fermi 架构特性(例如新的透明缓存的配置)之间的重要关系。我们提出了一种基于洞察的调优技术方法,提供了理解复杂关系的线路,并轻松避免不良的调优设置。
While the correctness of an NVIDIA CUDA program is easy to achieve, exploiting the GPU capabilities to obtain the best performance possible is a task for CUDA experienced programmers. Typical code tuning strategies, like choosing an appropriate size and shape for the thread-blocks, programming a good coalescing, or maximize occupancy, are inter-dependent. Moreover, the choices are also dependent on the underlying architecture details, and the global-memory access pattern of the designed solution. For example, the size and shapes of threadblocks are usually chosen to facilitate encoding (e.g. square shapes), while maximizing the multiprocessors' occupancy. However, this simple choice does not usually provide the best performance results. In this paper we discuss important relations between the size and shapes of threadblocks, occupancy, global memory access patterns, and other Fermi architecture features, such as the configuration of the new transparent cache. We present an insight based approach to tuning techniques, providing lines to understand the complex relations, and to easily avoid bad tuning settings.