GPU-STREAM v2.0: Benchmarking the Achievable Memory Bandwidth of Many-Core Processors Across Diverse Parallel Programming Models

GPU-STREAM v2.0: Benchmarking the Achievable Memory Bandwidth of Many-Core Processors Across Diverse Parallel Programming Models
复制标题

GPU-STREAM v2.0:跨多种并行编程模型对多核处理器可实现的内存带宽进行基准测试

DOI:
--
复制
发表时间:
2016
期刊:
ISC Workshops
影响因子:
--
通讯作者:
Simon McIntosh
Simon McIntosh
中科院分区:
--
文献类型:
--
作者:
Tom Deakin;J. Price;Matt Martineau;Simon McIntosh

文献摘要

被引文献

相似文献

许多科学代码由内存带宽绑定内核组成 - 运行时的主导因素是将数据从内存加载到算术逻辑单元中的速度,然后将结果写回到内存的一个主要优势。但是,由于通用图形处理单元(GPGPU)和英特尔Xeon Phi的重点是提供比传统CPU体系结构增加的记忆带宽。与CPU一样,这种峰值内存带宽通常在实践中是无法实现的,因此需要基准测量实用的上限对预期性能。
Many scientific codes consist of memory bandwidth bound kernels — the dominating factor of the runtime is the speed at which data can be loaded from memory into the Arithmetic Logic Units, before results are written back to memory. One major advantage of many-core devices such as General Purpose Graphics Processing Units (GPGPUs) and the Intel Xeon Phi is their focus on providing increased memory bandwidth over traditional CPU architectures. However, as with CPUs, this peak memory bandwidth is usually unachievable in practice and so benchmarks are required to measure a practical upper bound on expected performance.