Automated Instruction Stream Throughput Prediction for Intel and AMD Microarchitectures

Automated Instruction Stream Throughput Prediction for Intel and AMD Microarchitectures
复制标题

Intel 和 AMD 微架构的自动指令流吞吐量预测

DOI:
10.1109/pmbs.2018.8641578
复制
发表时间:
2018
期刊:
2018 IEEE/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS)
影响因子:
--
通讯作者:
G. Wellein
G. Wellein
中科院分区:
--
文献类型:
--
作者:
Jan Laukemann;Julian Hammer;Johannes Hofmann;G. Hager;G. Wellein

文献摘要

被引文献

相似文献

指令流的调度和执行的准确预测是预测在阶外处理器体系结构上预测吞吐量结合循环内核的核心性能行为的必要先决条件。这样的预测是分析性能模型的必不可少的组成部分,例如车顶线和执行 - 速度记忆(ECM)模型,并可以深入了解硬件体系结构与环路代码之间的性能相关的相互作用。我们提出了开源体系结构代码分析仪(OSACA),这是一种静态分析工具,用于预测在无限的第一级缓存和完美的doffer of阶外调度的假设下包含X86指令的顺序循环的执行时间。我们展示了从可用文档和半自动基准测试中构建机器模型的过程,并将其用于最新的Intel Skylake和AMD Zen Micro-Architectures。为了验证构造的模型,我们将它们应用于几个组装内核,并将运行时预测与实际测量结果进行比较。最后,我们对如何将方法推广到新体系结构进行展望。
An accurate prediction of scheduling and execution of instruction streams is a necessary prerequisite for predicting the in-core performance behavior of throughput-bound loop kernels on out-of-order processor architectures. Such predictions are an indispensable component of analytical performance models, such as the Roofline and the Execution-Cache-Memory (ECM) model, and allow a deep understanding of the performance-relevant interactions between hardware architecture and loop code. We present the Open Source Architecture Code Analyzer (OSACA), a static analysis tool for predicting the execution time of sequential loops comprising x86 instructions under the assumption of an infinite first-level cache and perfect out-of-order scheduling. We show the process of building a machine model from available documentation and semi-automatic benchmarking, and carry it out for the latest Intel Skylake and AMD Zen micro-architectures. To validate the constructed models, we apply them to several assembly kernels and compare runtime predictions with actual measurements. Finally we give an outlook on how the method may be generalized to new architectures.