A Performance Analysis of Vector Length Agnostic Code

A Performance Analysis of Vector Length Agnostic Code
复制标题

矢量长度不可知代码的性能分析

DOI:
--
复制
发表时间:
2019
期刊:
International Symposium on High Performance Computing Systems and Applications
影响因子:
--
通讯作者:
B. Juurlink
B. Juurlink
中科院分区:
--
文献类型:
--
作者:
Angela Pohl;Mirko Greese;Biagio Cosenza;B. Juurlink

文献摘要

被引文献

相似文献

向量扩展是近年来应用数据并行的一种流行,最常用的扩展名在矢量长度和矢量指令的数量上一直在增长。向矢量长度不可知论(VLA)架构已针对未来的ARM和RISC-V处理器进行,这些架构与这些体系结构相关。因此,目标硬件平台的矢量长度可以调整为通用矢量长度,以了解VLA代码的性能与矢量长度特定代码相比,我们分析了ARM SVE架构代码生成的当前功能。我们的实验表明,VLA代码大约达到了矢量长度特定代码的90%,即由于指令的全局预测,推断出10%的开销。我们表明,由于内存需求较高,代码性能并没有按比例增加,矢量长度的增加。
Vector extensions are a popular mean to exploit data parallelism in applications. Over recent years, the most commonly used extensions have been growing in vector length and amount of vector instructions. However, code portability remains a problem when speaking about a compute continuum. Hence, vector length agnostic (VLA) architectures have been proposed for the future generations of ARM and RISC-V processors. With these architectures, code is vectorized independently of the vector length of the target hardware platform. It is therefore possible to tune software to a generic vector length. To understand the performance impact of VLA code compared to vector length specific code, we analyze the current capabilities of code generation for ARM’s SVE architecture. Our experiments show that VLA code reaches about 90% of the performance of vector length specific code, i.e. a 10% overhead is inferred due to global predication of instructions. Furthermore, we show that code performance is not increasing proportionally with increasing vector lengths due to the higher memory demands.