Automatic Throughput and Critical Path Analysis of x86 and ARM Assembly Kernels

Automatic Throughput and Critical Path Analysis of x86 and ARM Assembly Kernels
复制标题

x86 和 ARM 汇编内核的自动吞吐量和关键路径分析

DOI:
10.1109/pmbs49563.2019.00006
复制
发表时间:
2019
期刊:
2019 IEEE/ACM Performance Modeling, Benchmarking and Simulation of High Performance Computer Systems (PMBS)
影响因子:
--
通讯作者:
G. Wellein
G. Wellein
中科院分区:
--
文献类型:
--
作者:
Jan Laukemann;Julian Hammer;G. Hager;G. Wellein

文献摘要

被引文献

相似文献

乱序架构上循环内核运行时的有用模型需要分析指令及其依赖关系的核内性能行为。当指令吞吐量预测为内核运行时设置下限时,关键路径定义上限。这种预测是分析的重要组成部分(即,白盒)性能模型,如Roofline和执行高速缓存存储器(ECM)模型。它们可以更好地理解硬件架构和循环代码之间与性能相关的交互。开源架构代码分析器(OSACA)是一个静态分析工具,用于预测顺序循环的执行时间。它以前只支持x86(Intel和AMD)架构和简单、乐观的全吞吐量执行。我们对OSACA进行了大量扩展,以支持ARM指令和关键路径预测,包括检测循环携带的依赖关系,这使其成为一个通用的跨架构建模工具。我们基于可用文档和半自动基准测试中的机器模型,展示了英特尔Cascade Lake、AMD Zen和马尔维尔ThunderX2微架构上代码的运行时预测。将预测结果与实际测量结果进行了比较。
Useful models of loop kernel runtimes on out-of-order architectures require an analysis of the in-core performance behavior of instructions and their dependencies. While an instruction throughput prediction sets a lower bound to the kernel runtime, the critical path defines an upper bound. Such predictions are an essential part of analytic (i.e., white-box) performance models like the Roofline and Execution-Cache-Memory (ECM) models. They enable a better understanding of the performance-relevant interactions between hardware architecture and loop code. The Open Source Architecture Code Analyzer (OSACA) is a static analysis tool for predicting the execution time of sequential loops. It previously supported only x86 (Intel and AMD) architectures and simple, optimistic full-throughput execution. We have heavily extended OSACA to support ARM instructions and critical path prediction including the detection of loop-carried dependencies, which turns it into a versatile cross-architecture modeling tool. We show runtime predictions for code on Intel Cascade Lake, AMD Zen, and Marvell ThunderX2 micro-architectures based on machine models from available documentation and semi-automatic benchmarking. The predictions are compared with actual measurements.