Benchmarking the achievable memory bandwidth of many-core processors across diverse parallel
Benchmarking the achievable memory bandwidth of many-core processors across diverse parallel
复制标题
对跨不同并行的多核处理器可实现的内存带宽进行基准测试
DOI:
--
复制
发表时间:
--
期刊:
影响因子:
--
通讯作者:
Simon McIntosh
中科院分区:
文献类型:
--
作者:
Tom Deakin;J. Price;Matt Martineau;Simon McIntosh
. Many scientific codes consist of memory bandwidth bound kernels — the dominating factor of the runtime is the speed at which data can be loaded from memory into the Arithmetic Logic Units, before results are written back to memory. One major advantage of many-core devices such as General Purpose Graphics Processing Units (GPGPUs) and the Intel Xeon Phi is their focus on providing increased memory bandwidth over traditional CPU architectures. However, as with CPUs, this peak memory bandwidth is usually unachievable in practice and so benchmarks are required to measure a practical upper bound on expected performance. The choice of one programming model over another should ideally not limit the performance that can be achieved on a device. GPU-STREAM has been updated to incorporate a wide variety of the latest parallel programming models, all implementing the same parallel scheme. As such this tool can be used as a kind of Rosetta Stone which provides both a cross-platform and cross-programming model array of results of achievable memory bandwidth.