Accelerating Level 2 BLAS Based on ARM SVE

Accelerating Level 2 BLAS Based on ARM SVE
复制标题

基于ARM SVE加速2级BLAS

DOI:
10.1109/aemcse51986.2021.00208
复制
发表时间:
2021
期刊:
2021 4th International Conference on Advanced Electronic Materials, Computers and Software Engineering (AEMCSE)
影响因子:
--
通讯作者:
Junjie Su
Junjie Su
中科院分区:
--
文献类型:
--
作者:
Xiuwen Wan;Naijie Gu;Junjie Su

文献摘要

被引文献

相似文献

Scalable Vector Extension(SVE)是ARM最近发布的一种特殊的向量指令集架构,是针对ARMv8-A架构的A64指令集的向量扩展。 SVE引入了许多重要的特性,不仅充分利用了宽向量的使用,而且还实现了高效的向量化函数,从而实现了高性能。本文描述了如何应用 SVE 的主要功能来加速 2 级 BLAS 例程。我们在ARM指令模拟器(ARMIE)和Gem5模拟器上进行了仿真。结果表明,与 NEON 实现相比,我们的 SVE 实现将执行指令数量减少了 90%,并实现了超过 17 倍的加速。
Scalable Vector Extension (SVE) is a special vector instruction set architecture recently released by ARM, which is a vector extension for A64 instruction set of ARMv8-A architecture. SVE introduces many significant features that not only take full advantage of the use of wide vector, but also enable efficient vectorization functions which can realize high performance. This paper described how to apply the main features of SVE to accelerate the Level 2 BLAS routines. We carried out the simulation on ARM Instruction Emulator (ARMIE) and Gem5 simulator. The results demonstrated that our SVE implementation reduced the number of executed instructions by 90% and achieved over 17x speedup compared with the NEON implementation.