Nested data-parallelism on the gpu

Nested data-parallelism on the gpu
复制标题

GPU 上的嵌套数据并行性

DOI:
10.1145/2364527.2364563
复制
发表时间:
2012
期刊:
Proceedings of the 17th ACM SIGPLAN international conference on Functional programming
影响因子:
--
通讯作者:
John H. Reppy
John H. Reppy
中科院分区:
--
文献类型:
--
作者:
Lars Bergstrom;John H. Reppy

文献摘要

被引文献

相似文献

图形处理单元(GPU)提供的内存带宽和算术性能远远大于CPU上可用的性能,但是由于其单一指令 - 元素 - data(SIMD)体系结构,它们很难编程。到目前为止,大多数已移植到GPU的程序都使用传统的数据级并行性,仅执行均匀操作的操作。 NESL是一种一阶功能语言,旨在允许程序员为宽矢量并行计算机编写不规则平行程序(例如平行分隔和串联算法)。本文介绍了我们的NESL实施港口以在GPU上运行的港口,并提供了嵌套数据并行性(NDP)的经验证据,在GPU上嵌套了大大优于基于CPU的实现和匹配或匹配或击败仅支持平行平行性的较新的GPU语言。尽管我们的性能与手工调整的CUDA程序的性能不符,但我们认为NESL的符号简洁性值得损失。这项工作提供了直接支持GPU上NDP的第一语言实现。
Graphics processing units (GPUs) provide both memory bandwidth and arithmetic performance far greater than that available on CPUs but, because of their Single-Instruction-Multiple-Data (SIMD) architecture, they are hard to program. Most of the programs ported to GPUs thus far use traditional data-level parallelism, performing only operations that operate uniformly over vectors. NESL is a first-order functional language that was designed to allow programmers to write irregular-parallel programs - such as parallel divide-and-conquer algorithms - for wide-vector parallel computers. This paper presents our port of the NESL implementation to work on GPUs and provides empirical evidence that nested data-parallelism (NDP) on GPUs significantly outperforms CPU-based implementations and matches or beats newer GPU languages that support only flat parallelism. While our performance does not match that of hand-tuned CUDA programs, we argue that the notational conciseness of NESL is worth the loss in performance. This work provides the first language implementation that directly supports NDP on a GPU.