Evaluating attainable memory bandwidth of parallel programming models via BabelStream

Evaluating attainable memory bandwidth of parallel programming models via BabelStream
复制标题

通过 BabelStream 评估并行编程模型可达到的内存带宽

DOI:
10.1504/ijcse.2017.10011352
复制
发表时间:
2018
期刊:
Int. J. Comput. Sci. Eng.
影响因子:
--
通讯作者:
Simon McIntosh
Simon McIntosh
中科院分区:
--
文献类型:
--
作者:
Tom Deakin;J. Price;Matt Martineau;Simon McIntosh

文献摘要

被引文献

相似文献

许多科学代码包括内存带宽绑定内核。多核设备(例如通用图形处理单元(GPGPU)和Intel Xeon Phi)等多核设备的一个主要优点是,他们的重点是提供比传统CPU体系结构增加的内存带宽。峰值内存带宽通常在实践中是无法实现的,因此需要基准测量实用的上限对预期性能。我们使用DOT产品内核增强标准流核,以调查大型阵列中简单减少操作的性能。理想情况下,编程模型的选择不应限制设备上可实现的性能。 BabelStream(正式的GPU-stream)已更新,以结合各种最新的并行编程模型,所有这些模型都实现了相同的并行方案。因此,该工具可以用作一种罗塞塔石,既可以提供可实现的内存带宽结果的跨平台和交叉编程模型阵列。
Many scientific codes consist of memory bandwidth bound kernels. One major advantage of many-core devices such as general purpose graphics processing units (GPGPUs) and the Intel Xeon Phi is their focus on providing increased memory bandwidth over traditional CPU architectures. Peak memory bandwidth is usually unachievable in practice and so benchmarks are required to measure a practical upper bound on expected performance. We augment the standard STREAM kernels with a dot product kernel to investigate the performance of simple reduction operations on large arrays. The choice of programming model should ideally not limit the achievable performance on a device. BabelStream (formally GPU-STREAM) has been updated to incorporate a wide variety of the latest parallel programming models, all implementing the same parallel scheme. As such this tool can be used as a kind of Rosetta Stone which provides both a cross-platform and cross-programming model array of results of achievable memory bandwidth.