Automatic loop kernel analysis and performance modeling with Kerncraft

Automatic loop kernel analysis and performance modeling with Kerncraft
复制标题

使用 Kerncraft 自动循环内核分析和性能建模

DOI:
--
复制
发表时间:
2015
期刊:
International Workshop on Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems
影响因子:
--
通讯作者:
G. Wellein
G. Wellein
中科院分区:
--
文献类型:
--
作者:
Julian Hammer;G. Hager;Jan Eitzinger;G. Wellein

文献摘要

被引文献

相似文献

分析性能模型对于理解循环内核的性能特性是必不可少的,循环内核消耗了计算科学中CPU周期的主要部分。从经过验证的性能模型开始,可以推断出相关的硬件瓶颈和有希望的优化机会。不幸的是,即使对于有经验的开发人员来说,分析性能建模也往往是乏味的,因为它需要深入了解硬件及其如何与软件交互。我们提出了“Kerncraft”工具,它简化了流内核和模板循环嵌套的分析性能模型的构建。从循环源代码、问题大小和底层硬件描述开始,Kerncraft可以使用Roofline或Execution-Cache-Memory(ECM)模型来理想地预测多核处理器上循环的单核性能和扩展行为。我们描述了Kerncraft的操作原理及其功能和局限性,并展示了如何通过加速分析建模快速获得见解。
Analytic performance models are essential for understanding the performance characteristics of loop kernels, which consume a major part of CPU cycles in computational science. Starting from a validated performance model one can infer the relevant hardware bottlenecks and promising optimization opportunities. Unfortunately, analytic performance modeling is often tedious even for experienced developers since it requires in-depth knowledge about the hardware and how it interacts with the software. We present the "Kerncraft" tool, which eases the construction of analytic performance models for streaming kernels and stencil loop nests. Starting from the loop source code, the problem size, and a description of the underlying hardware, Kerncraft can ideally predict the single-core performance and scaling behavior of loops on multicore processors using the Roofline or the Execution-Cache-Memory (ECM) model. We describe the operating principles of Kerncraft with its capabilities and limitations, and we show how it may be used to quickly gain insights by accelerated analytic modeling.