Designing Efficient Index-Digit Algorithms for CUDA GPU Architectures

Designing Efficient Index-Digit Algorithms for CUDA GPU Architectures
复制标题

为 CUDA GPU 架构设计高效的索引数字算法

DOI:
10.1109/tpds.2015.2450718
复制
发表时间:
2016
影响因子:
5.3
通讯作者:
R. Doallo
R. Doallo
中科院分区:
计算机科学2区
文献类型:
--
作者:
J. Lobeiras;M. Amor;R. Doallo

文献摘要

被引文献

相似文献

现代图形处理单元(GPU)以相对较低的成本提供很高的计算能力。但是,即使对于经验丰富的程序员,为GPU设计有效的算法通常需要额外的时间和精力。在这项工作中,我们提出了一种调整方法,该方法允许设计启用索引数字算法的CUDA GPU架构,即,可以将数据移动描述为包含数据元素索引的数字的排列的算法。该方法基于确定为GPU资源分析和运算符的两个阶段,用于FFT和Tridia-Gonal Systems求解器算法,分析性能特征和最适当的解决方案。由此产生的实施是紧凑的,胜过其他著名和常用的最先进的库,比NVIDIA的复杂袖口提高了19.2%,而NVIDIA'Scudpp的实际数据是三大数据,而实际数据则超过3000%。系统。
Modern graphics processing units (GPUs) offer very high computing power at relatively low cost. Nevertheless, designing efficient algorithms for the GPUs normally requires additional time and effort, even for experienced programmers. In this work we present a tuning methodology that allows the design for CUDA-enabled GPU architectures of index-digit algorithms, that is, algorithms where the data movement can be described as the permutations of the digits comprising the indices of the data elements. This methodology, based on two-stages identified as GPU resource analysis and operators string manipulation, is applied to FFT and tridiagonal systems solver algorithms, analyzing the performance features and the most adequate solutions. The resulting implementation is compact and outperforms other well-known and commonly used state-of-the-art libraries, with an improvement of up to 19.2 percent over NVIDIA's complex CUFFT , and more than 3000 percent over the NVIDIA'sCUDPP for real data tridiagonal systems.