Benchmarking the achievable memory bandwidth of many-core processors across diverse parallel

Benchmarking the achievable memory bandwidth of many-core processors across diverse parallel
复制标题

对跨不同并行的多核处理器可实现的内存带宽进行基准测试

DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
Simon McIntosh
Simon McIntosh
中科院分区:
--
文献类型:
--
作者:
Tom Deakin;J. Price;Matt Martineau;Simon McIntosh

文献摘要

被引文献

相似文献

。许多科学代码都由内存带宽限制的内核组成 - 运行时的主导因素是在结果写回内存之前将数据从内存加载到算术逻辑单元的速度。通用图形处理单元 (GPGPU) 和 Intel Xeon Phi 等多核设备的一大优势是它们专注于提供比传统 CPU 架构更高的内存带宽。然而,与 CPU 一样,这种峰值内存带宽在实践中通常无法实现,因此需要基准测试来衡量预期性能的实际上限。理想情况下,选择一种编程模型而不是另一种编程模型不应限制设备上可以实现的性能。 GPU-STREAM 已更新,包含各种最新的并行编程模型,所有模型均实现相同的并行方案。因此,该工具可以用作一种 Rosetta Stone,它提供了可实现的内存带宽结果的跨平台和跨编程模型数组。
. Many scientific codes consist of memory bandwidth bound kernels — the dominating factor of the runtime is the speed at which data can be loaded from memory into the Arithmetic Logic Units, before results are written back to memory. One major advantage of many-core devices such as General Purpose Graphics Processing Units (GPGPUs) and the Intel Xeon Phi is their focus on providing increased memory bandwidth over traditional CPU architectures. However, as with CPUs, this peak memory bandwidth is usually unachievable in practice and so benchmarks are required to measure a practical upper bound on expected performance. The choice of one programming model over another should ideally not limit the performance that can be achieved on a device. GPU-STREAM has been updated to incorporate a wide variety of the latest parallel programming models, all implementing the same parallel scheme. As such this tool can be used as a kind of Rosetta Stone which provides both a cross-platform and cross-programming model array of results of achievable memory bandwidth.